หุ่นยนต์และ AI ทางกายภาพ

Motional เปิดตัวชุดข้อมูล nuReasoning และเปิดการแข่งขัน ECCV

mm
เพิ่ม Unite.AI ลงในแหล่งข้อมูลที่คุณต้องการบน Google

Motional ประกาศการเปิดตัว nuReasoning เมื่อวันที่ 8 กันยายน 2026 โดยอธิบายว่าเป็นชุดข้อมูลเปิดแบบสถานการณ์หางยาวที่มุ่งเน้นการให้เหตุผลเป็นครั้งแรกของโลกสำหรับยานยนต์อัตโนมัติ และเปิดการแข่งขันวิจัยคู่ขนานที่การประชุม European Conference on Computer Vision ในสวีเดน ชุดข้อมูลนี้สร้างร่วมกับ UCLA Mobility Lab และผู้กำกับของห้องปฏิบัติการ ศาสตราจารย์ Jiaqi Ma.

The ประกาศของ Motional positions nuReasoning as training material for end-to-end autonomous systems, targeting the rare edge cases where perception alone is insufficient. Vision-Language-Action models combine scene recognition with decisions about how to act, and Motional states that such models need training data that conveys the logic behind a driving action rather than the correct action alone. The dataset is designed to teach models spatial relationships, driving-decision reasoning, risk anticipation, and consideration of alternative outcomes.

โครงสร้างชุดข้อมูล

nuReasoning มีสถานการณ์หางยาวจำนวน 20,000 รายการ โดยแต่ละรายการเป็นคลิปวิดีโอที่ยาวอย่างน้อย 20 วินาทีและฝังคำอธิบายการให้เหตุผลที่ได้รับการตรวจสอบโดยมนุษย์ เหตุการณ์เหล่านี้สกัดจากข้อมูลฝูงเรือของ Motional ที่เก็บรวบรวมใน Las Vegas, Pittsburgh, Los Angeles, Boston และ Singapore และบริษัทรายงานว่ามีกรณีขอบที่ต้องการการให้เหตุผลมากกว่า 105 ชั่วโมง รวมถึงกิจกรรมคนเดินเท้าแปลกประหลาด โซนทำงานและการก่อสร้างถนนในเวลากลางคืน การข้ามสัตว์ และสถานการณ์ที่มองเห็นได้จำกัด ชุดคำอธิบายประกอบด้วยคำอธิบายการให้เหตุผลที่ตรวจสอบโดยมนุษย์ทั้งหมด 247,000 รายการ แบ่งเป็นสามประเภท: Spatial Reasoning, Decision Reasoning, และ Counterfactual Reasoning พร้อมกับข้อมูลเซ็นเซอร์หลายโหมดสำหรับการแสดงฉากสามมิติ

The official บันทึก nuReasoning อย่างเป็นทางการ provides additional structure. Clips are approximately 20 seconds long and sampled at 10 Hz, with 17,000 training clips, 2,000 validation clips, and 1,000 private test clips. Sensor data includes synchronized multi-view cameras, LiDAR for supported clips, calibration and ego state, 3D object annotations, vectorized HD maps, traffic lights, route paths, and navigation commands. To select the long-tail material, an AI evaluator scored more than 100,000 pre-filtered 30-second candidate clips drawn from over 10,000 hours of driving, followed by human verification and keyframe selection.

Spatial Reasoning annotations cover object-centric 2D-3D understanding, map and topology context, relative position, velocity, future motion, and interactions with the ego vehicle, annotated at 1 Hz. Decision Reasoning captures the scene description, critical components, longitudinal and lateral meta-actions, and a causal trace explaining the selected behavior. Counterfactual Reasoning evaluates alternative ego actions against safety-critical outcomes and risk levels labeled safe, suboptimal, or unsafe; decision and counterfactual annotations are recorded at 0.2 Hz. The record also states that the dataset applies privacy-preserving processing, including anonymization of human faces and license plates.

Motional has integrated its proprietary Omnitag search engine, which allows researchers to query distributions by scenario type, difficulty level, and location, and to run natural-language semantic searches for tactical interactions. One annotated scenario described by the company shows a nighttime construction zone where the vehicle stopped: the annotation explains that the stop was correct because of a small animal crossing ahead, and that alternate routes were rejected due to construction barriers and the animal’s unknown speed and direction.

The company reports that a nuReasoning miniset released earlier this year has been downloaded more than 50,000 times. The public release covers the training and validation splits on Hugging Face, while the 1,000 private test clips are withheld for challenge evaluation. An open-source devkit on GitHub provides data processing for VLA training, reasoning and planning targets, baseline models, and evaluation metrics.

การแข่งขัน nuReasoning

Motional and the UCLA Mobility Lab are hosting the nuReasoning Challenge, launched at ECCV. The competition is structured as a joint evaluation of planning and reasoning on the 1,000 private-test scenarios. The main planning task gives participants 10 seconds of driving context — multi-view images, HD map information, navigation command, and ego state — and requires a predicted ego trajectory for the next five seconds. The auxiliary reasoning task uses the same context for visual question answering spanning geometry, motion, decision, and counterfactual reasoning, with multiple-choice, numerical, coordinate, and trajectory answer formats.

The final challenge score combines 75% planning score and 25% reasoning score, and the total prize pool is $10,000. Registration is open, and winners will be announced in ธันวาคม 2026 at the Conference on Neural Information Processing Systems.

“To safely expand AV fleets, autonomous driving systems must react to rare, chaotic edge cases with the same assured logic as an experienced human driver,” said Phil Michel, Motional’s Senior Vice President of Autonomy and AI. He said the company is prioritizing transparent AI so real-time decision-making and risk assessment can be clearly evaluated, and that making nuReasoning openly available offers a shared foundation for solving edge cases.

พื้นฐานการวิจัยและการเข้าถึง

The accompanying paper, titled “nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving,” was submitted to arXiv on 29 พฤษภาคม 2026, with authors affiliated with UCLA and Motional and Tony Xuewei Qi listed as project lead. In the paper, the authors report that fine-tuning vision-language models on nuReasoning substantially improves driving-specific question answering, and that incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. The paper’s planning model, nuVLA, combines a vision-language backbone with trajectory generation.

nuReasoning extends Motional’s open-data line that began with nuScenes in 2019, which the company identifies as the industry’s first multi-modal public AV dataset, and continued through nuImages, Panoptic nuScenes, and nuPlan. The datasets are available under commercial licenses or free academic use subject to Motional’s non-commercial terms and conditions.

เอลารา นิกซ์ เป็นนักวิเคราะห์ที่สร้างโดย AI ที่ Unite.AI โดยครอบคลุมเรื่องปัญญาประดิษฐ์ในด้านการขนส่ง ระบบการเดินทาง และเทคโนโลยีแบบอัตโนมัติ ผลงานของเธอเน้นไปที่วิธีการที่ AI กำลังเปลี่ยนแปลงยานพาหนะที่ขับเคลื่อนด้วยตนเอง ระบบการบิน รางและระบบขนส่งสาธารณะ - ที่ที่ที่ความน่าเชื่อถือ ความปลอดภัย และการนำไปใช้ในโลกแห่งความเป็นจริงมีความสำคัญไม่แพ้ความก้าวหน้า

ด้วยมุมมองทางเทคนิคและที่มองไปข้างหน้า เอลารา ตรวจสอบความก้าวหน้าในระบบการรับรู้ ระบบอัตโนมัติ การจำลองและการตรวจสอบความปลอดภัยทั่วทั้งพื้นที่ ทั้งทางบก ทางอากาศ และระบบการเดินทางในเมือง เธอสนใจเป็นพิเศษเกี่ยวกับวิธีการที่ระบบอัตโนมัติที่ขับเคลื่อนด้วย AI ถูกทดสอบ ถูกควบคุม และรวมเข้ากับโครงสร้างพื้นฐานที่มีอยู่แล้ว และวิธีการที่ระบบเหล่านี้สร้างสมดุลระหว่างความได้เปรียบด้านประสิทธิภาพกับความไว้วางใจของสาธารณชนและการจัดการความเสี่ยง

บทความที่เขียนโดย Elara Nix สร้างโดย AI และได้รับการตรวจสอบโดยทีมบรรณาธิการของ Unite.AI เพื่อให้แน่ใจถึงความถูกต้อง ความชัดเจน และการรายงานที่มีความรับผิดชอบเกี่ยวกับบทบาทที่เปลี่ยนแปลงของ AI ในด้านการขนส่งและระบบอัตโนมัติ