Research ·
Making Progress Measurable: A Common Task Framework for Autonomous Driving with AlpaSim
Our previous post introduced the research vehicle we will use for real-world testing, helping us understand and account for hardware-related challenges such as sensor noise and calibration. However, systematic progress also requires experiments that other researchers can repeat. This is the motivation behind the common task framework[1]: shared development data, a defined evaluation protocol, held-out test data, and a public leaderboard. Researchers evaluate their methods under a set of common rules on examples withheld from development and openly share their findings. This approach has helped organize and accelerate research in vision, speech, and language since the 1980s. We now aim to bring that same discipline to the development of frontier models for real-world end-to-end autonomous driving.
Together with NVIDIA’s ASPIRE Group and HKU’s OpenDriveLab, we are organizing the AlpaSim End-to-End Closed-Loop Challenge[2]. The challenge is the first to bring together photorealistic neural simulation[3], accessible baselines, and large evaluation suites in a shared framework for comparing end-to-end driving policies. Our aim is to make frontier autonomous driving research accessible to a broader community.
Lightweight simulation with NAVSIM
A recorded drive shows one way to navigate a scene, but other choices can be safe too. Early attempts to establish an accessible common task for autonomous driving evaluated a method’s planned path by its distance from the human driver’s recorded path: a comparison that can reward the wrong behavior[4]. Try the controls in Figure 1: changing speed while staying in the lane can produce a larger average displacement error (ADE) than swerving off the road.
NAVSIM[5] was our attempt to address this problem by simulating a planned path for a few seconds (without replanning) and measuring safety, comfort, and progress. We used NAVSIM to test policies on large and carefully curated real-world driving datasets, without relying on the ADE metric. The simplified NAVSIM score below illustrates the benefit of this approach: it penalizes leaving the road or entering oncoming traffic, even when the path’s endpoint stays close to the recording.
Change speed
✓ Compliant
Swerve while maintaining speed
✕ Drivable area compliance
NAVSIM now provides a practical first step for evaluating new ideas in autonomous driving. Since 2024, we have organized three successful NAVSIM challenges. The inaugural 2024 challenge[6] attracted 143 teams and 463 submissions, showing how much research a shared evaluation setup can support.
We have also improved NAVSIM over time, introducing pseudo-simulation[7] to probe how a policy recovers from deviations. Before evaluation, it uses 3D Gaussian Splatting[8] to render additional observations at different positions, headings, and speeds. During evaluation, these observations receive weights according to how closely they match states the policy is likely to reach. This lets the benchmark test recovery beyond the recorded drive.
However, the common task framework requires us to revisit the choice of task once community progress begins to saturate. Unfortunately, NAVSIM does not account for important deployment constraints. A policy can score well while far exceeding a typical vehicle’s GPU memory budget or taking too long to produce a plan. Incorporating significant methodological complexity is now common on the benchmark, even for marginal gains. Therefore, the AlpaSim challenge moves to a closed-loop simulation setup with explicit compute and timing limits.
Closed-loop simulation at scale with AlpaSim
The AlpaSim challenge uses an open, modular simulator for closed-loop evaluation on two independent tracks: the Physical AI AV dataset[9], with 1,700 hours of driving data, and nuPlan[10], with 120 hours of publicly released sensor data. A common simulator lets researchers develop policies across different geographies and vehicle configurations without maintaining separate evaluation setups. This enables the exploration of new research questions beyond NAVSIM, which only supported the nuPlan dataset. Our developer kit for the AlpaSim challenge enables joint training on both Physical AI AV and nuPlan, making it easier to draw on their combined scale and diversity. Figure 2 shows a baseline policy trained with the developer kit driving on both challenge tracks.
Like NAVSIM, AlpaSim penalizes policies for collisions, off-road driving, and failure to make progress, without using traditional displacement errors. Unlike NAVSIM’s pseudo-simulation, which weights pre-rendered observations, AlpaSim runs a full closed-loop simulation: each action changes the simulated vehicle’s state, and newly rendered observations feed back into the policy’s next decision. This lets evaluation follow the consequences of successive decisions and test recovery as a drive unfolds. Furthermore, the challenge makes computational efficiency part of the task. Policies must fit within 16 GiB of GPU memory and target a planning frequency of 10 Hz, producing a new plan every 0.1 seconds. This brings evaluation closer to the demands of running a policy on a real vehicle.
Closed-loop simulation requires substantial compute, which can put large-scale evaluation out of reach for smaller teams. To lower this barrier, participants in the AlpaSim challenge submit a containerized policy, and the organizers provide the compute and infrastructure for official evaluation. Our test suite covers thousands of private held-out scenes across Denmark, France, Germany, Japan, Singapore, South Korea, Spain, Sweden, the United Kingdom, and the United States.
To keep evaluation fair, each team receives the same fixed number of official submissions. We also provide smaller, standardized validation splits so teams can run rigorous ablation studies on local compute before using their official leaderboard submissions.
Towards multitask benchmarking with Item Response Theory
As driving policies become more general, we need to evaluate them across multiple datasets and sensor configurations. This raises a question that also appears in multitask benchmarks for large language models: how should results from different tasks be combined into a meaningful assessment of capability? Simply averaging scores makes implicit choices about the importance and difficulty of each task. A model’s ranking can then depend on the benchmark’s composition[11], and improvements in the aggregate can hide weaknesses in individual skills. Frontier driving models will face the same problem as evaluation expands across environments.
We are exploring Item Response Theory (IRT) as an initial approach to this problem. Originally developed for educational testing[12], IRT treats each policy as a student and each driving scenario as an exam question. It estimates policy ability and scenario difficulty from the pattern of successes and failures across policies. This provides a way to account for the difficulty of scenarios in which a policy succeeds when assigning scores and to identify scenarios that best distinguish policies. IRT also estimates uncertainty in policy ability, allowing us to show a range of plausible ranks for each policy. We applied IRT to existing NAVSIM submissions as an illustrative example, displaying policy ability and route difficulty rankings with rank intervals.
The current AlpaSim challenge will be our first attempt to put this approach into practice on a live leaderboard. IRT will not be used to combine the two tracks, but rather to score the policies and routes independently in each track. For policies with overlapping rank intervals, we will use kilometers driven per at-fault infraction as a ranking metric to break the tie. We aim to use what we learn from the challenge to refine our approach and continue investigating solutions for curating and expanding the AlpaSim leaderboard tracks.
Conclusion
Tackling the hardest problems in autonomous driving requires a shared research effort. The AlpaSim challenge is our next step towards making this possible. There is still much to learn about how best to measure driving capability, and we expect the benchmark to evolve as its limitations become clear. We hope it can provide a common task for the end-to-end driving community to gather around, leading to the discovery of new techniques to improve driving policy safety and efficiency.
The challenge opened on June 15, 2026. Under the planned timeline, upon completion of the ongoing maintenance, the rules and submission formats freeze on September 15. The public leaderboard closes on October 31, when final submissions and technical reports are due. You can follow the live results in Figure 3:
References
- M. Liberman. “Reproducible Research and the Common Task Method.” Simons Foundation, 2015.
- NVIDIA. “AlpaSim End-to-End Closed-Loop Challenge.” Hugging Face, 2026.
- Y. Zhang, K. Tóthová, Z. Wang, K. Yin, H. Turki, R. de Lutio, Y.-Y. Chang, O. Litany, S. Fidler, Z. Gojcic. “DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer.” arXiv:2602.24096, 2026.
- B. Jaeger, K. Chitta, D. Dauner, K. Renz, A. Geiger. “Common Mistakes in Benchmarking Autonomous Driving.” 2024.
- D. Dauner et al. “NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking.” NeurIPS, 2024.
- OpenDriveLab. “End-to-End Driving at Scale.” Autonomous Grand Challenge, CVPR 2024.
- W. Cao et al. “Pseudo-Simulation for Autonomous Driving.” CoRL, 2025.
- B. Kerbl, G. Kopanas, T. Leimkühler, G. Drettakis. “3D Gaussian Splatting for Real-Time Radiance Field Rendering.” ACM Transactions on Graphics (SIGGRAPH), 2023.
- NVIDIA. “Physical AI Autonomous Vehicles Dataset.” Hugging Face.
- N. Karnchanachari et al. “Towards Learning-Based Planning: The nuPlan Benchmark for Real-World Autonomous Driving.” ICRA, 2024.
- G. Zhang, M. Hardt. “Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks.” ICML, 2024.
- A. Birnbaum. “Some Latent Trait Models and Their Use in Inferring an Examinee’s Ability.” In F. M. Lord & M. R. Novick (Eds.), Statistical Theories of Mental Test Scores. Addison-Wesley, Reading, MA, 1968, pp. 397–479.