Evaluating Spatiotemporal Reasoning in Driving Edge Cases

Lieqi Liu1*, Rui Gao1*, Jia-Chen Gu1, Wenbo Hu1, Zhaobin Mo2, Ahmadreza Moradipari2, Nejib Ammar2, Wei Wang1, Nanyun Peng1

1University of California, Los Angeles    2Toyota InfoTech Labs

*Equal contribution

Preprint, 2026

Overview of STRIDE: data, tasks, and comparison to prior driving VQA

Figure 1. Overview of STRIDE. Five-frame driving context, six task families, and a sharp drop versus prior driving VQA suites (GPT-6-Astra: 85.2% โ†’ 45.5% MCQ on nuScenes).

Introduction

Vision-language models look strong on existing driving VQA benchmarks, but those suites often test static recognition rather than the spatiotemporal reasoning needed in edge cases. STRIDE asks whether a model can track objects over time, judge driving significance, and plan ego motion from multi-frame observations.

The benchmark contains 2,350 questions over nuScenes (1,150) and WOD-E2E (1,200), spanning 39 templates in six families: spatial perception, spatial understanding, temporal memory, temporal extrapolation, trajectory prediction, and scene-context awareness. Formats include MCQ, open-ended questions (OEQ), and geometric waypoint regression. The strongest model we evaluated, GPT-6-Astra, reaches only 45.5% MCQ on nuScenes and 41.2% on Waymo โ€” well above random chance (~20%), but far from prior driving VQA numbers.

2,350questions
39templates ยท 6 families
5frames per question
45.5%best MCQ (nuScenes)

Leaderboard

Primary ranking key: MCQ accuracy. Open-ended (OEQ) and trajectory (ADE / mini-FDE) metrics are reported separately. Missing entries are shown as โ€”.

#ModelType MCQ โ†‘OEQ โ†‘ ADE โ†“mini-FDE โ†“
1 ๐Ÿฅ‡gpt-6-astraproprietary45.56.410.600.98
2 ๐ŸฅˆGPT-5.5proprietary39.36.680.650.97
3 ๐Ÿฅ‰gemini-3.6-flashproprietary38.96.500.691.01
4gpt-5.6-solproprietary38.55.650.690.99
5gemini-3.1-pro-previewproprietary37.84.531.421.84
6Qwen3.6 35B-A3BVLM31.65.327.425.45
7claude-sonnet-5proprietary29.46.420.871.30
8Qwen2.5-VL 3BVLM27.43.62โ€”โ€”
9Qwen3-VL 30B-A3B-ThinkingVLM26.85.457.123.42
10Qwen3-VL 8B-ThinkingVLM23.45.508.003.80
11Dolphinsexpert18.53.248.163.81
12Qwen3.6 27BVLM17.55.995.814.06
โ€”UniAD 2.0expert (TRJ-5)โ€”โ€”0.560.79
*Random chanceโ€”19.8โ€”โ€”โ€”

OEQ is a 0โ€“10 vision-judge score on open-ended answers. Trajectory errors are in meters (nuScenes vs Waymo scales differ). Machine-readable tables live in the repo. To submit a result, open an issue or PR with your *_score_report.json.

Benchmark

Each question uses a five-frame temporal window. nuScenes inputs are 2ร—3 surrounding-camera grids; Waymo uses CAM_FRONT. When a target object exists, it is marked with a red box on the query frame. Annotations ship in this repository and on Hugging Face; raw images are downloaded from nuScenes and Waymo Open Dataset.

Space

Spatial perception ยท SP

Observable object properties: relative position, distance, motion direction, lane occupancy.

Space

Spatial understanding ยท SU

Driving significance of spatial relations: drivable area, lane change, ego constraint, risk.

Time

Temporal memory ยท TM

How the current state formed: prior visibility, occlusion, reappearance, recent motion.

Time

Temporal extrapolation ยท TE

Likely evolution: future motion, occupancy, collision timing, subsequent events.

Context

Trajectory prediction ยท TRJ

Ego past/future motion via ranking, language justifications, and waypoint regression.

Context

Scene-context awareness ยท SC

Environment, road structure, and surrounding activity that frame individual objects.

Setup and evaluation instructions: github.com/PlusLabNLP/STRIDE.

Statistics

STRIDE covers challenging driving situations (pedestrians, intersections, construction, debris) across both source datasets, with object-centric questions spanning distance and bearing relative to the ego vehicle.

Scenario, task-family, and ego-behavior coverage of STRIDE

Figure 2. Dataset coverage: (a) scenario tags on nuScenes vs Waymo, (b) task-family mix, (c) future ego speed change by maneuver.

Spatial coverage heatmap of target objects

Figure 3. Spatial coverage of object-centric questions (bearing ร— distance).

Scenario coverage across task families

Figure 4. Scenario coverage across the six task families (multi-label clips).

Citation

@misc{liu2026stride,
  title        = {{STRIDE}: Evaluating Spatiotemporal Reasoning in Driving Edge Cases},
  author       = {Liu, Lieqi and Gao, Rui and Gu, Jia-Chen and Hu, Wenbo and Mo, Zhaobin
                  and Moradipari, Ahmadreza and Ammar, Nejib and Wang, Wei and Peng, Nanyun},
  year         = {2026},
  note         = {Preprint},
  howpublished = {\url{https://github.com/PlusLabNLP/STRIDE}}
}