NVIDIA’s Alpamayo 2 Super tackles the toughest gaps in today’s autonomous‑driving stacks: rare, multi‑agent events that traditional detection‑and‑prediction pipelines miss, the lack of clear causal explanations for safety validation, and the costly, slow process of labeling large fleets of video data. Teams also struggle with licensing uncertainty when trying to commercialize research models and with the heavy compute needed to run large vision‑language‑action systems in real‑time vehicles.
Alpamayo 2 Super offers a direct, practical answer. Released under the permissive OpenMDW‑1.1 license with Apache 2.0 source code, it can be fine‑tuned, redistributed, and used commercially from day one—no extra permissions required. The 34‑parameter model fuses a 32B Cosmos 3 Super Reasoner vision‑language backbone with a 2.3B diffusion action decoder, producing in a single forward pass: a 64‑waypoint trajectory (0.1‑6.4 s), a Chain‑of‑Causation text trace that links perception to action, a compact meta‑action label (yield, lane change, stop), reasoning auto‑labels for rapid fleet annotation, and grounded visual question answering.
Benchmarks show it leads the field—LingoQA Lingo‑Judge score 79.2, outperforming Qwen2.5‑VL 72B by +17.0, Gemini 2.5 Pro by +15.1, and GPT‑4o by +23.2. Closed‑loop AlpaSim scores 1.50 ± 0.13 on 910 nuanced scenarios, and open‑loop minADE₆ at 6.4 s is 0.911 m. Tested on a single H100 80GB GPU with ~72 GiB peak memory, the model can be distilled for lower‑power automotive hardware, turning months of annotation work into days and giving engineers a transparent, licensable foundation for safer, more reliable self‑driving systems.
#AI #Product #AutonomousDriving #VLA #OpenModel #NVIDIA