CRC 408 AgiMo · project C6

LLM Benchmark for AgiMo

Compare transport knowledge by task. Explore the leading results, then inspect the evidence.

Current benchmark · 15 tasks · 1088 questions

Who leads the most tasks

Score per task type: 0 = naive answer, 100 = exact (accuracy × 100 where there is no naive answer).

Gemma 4 26B

2of 15 tasks
best model
  • 2a Graphs82
  • 2b Local knowledge-31
  • 4 Mode choice-2

Llama 3.3 70B

3of 15 tasks
best model
  • 2a Graphs40
  • 2b Local knowledge-17
  • 4 Mode choice-4

Qwen3.8 27B (thinking off)

4of 15 tasks
best model
  • 2a Graphs55
  • 2b Local knowledge-13
  • 4 Mode choice-32

Qwen3.8 27B (thinking on)

6of 15 tasks
best model
  • 2a Graphs74
  • 2b Local knowledge-18
  • 4 Mode choice-11

Mobility planning lifecycle

Every benchmark task belongs to one stage of the planning cycle.

AgiMoPlanning cycle 12a2b345
1Setting goalsNot covered yet
2aBaseline analysis (graph theory)3 tasks
2bBaseline analysis (local knowledge)9 tasks
Station mode
District motorisation
Table information extraction
3Develop options1 task
Brainstorming scenarios
4Evaluate impact3 tasks
5Engage stakeholdersNot covered yet
MODEL × TASK

Compare the evidence.

Best result per task is highlighted. Scroll horizontally for all tasks.

Task details →
2a Graphs2b Local knowledge4 Mode choice
Traffic volumeStation modeDistrict motorisationTable information extractionMode choice prediction
Model / configurationReachabilityexact accuracy ↑ · 371 qMaximum flowexact accuracy ↑ · 58 qShortest pathexact accuracy ↑ · 64 qBerlin vehicle-kmaccuracy ↑ · 50 qDresden station volumeaccuracy ↑ · 50 qDresden heavy-vehicle shareMAE pp ↓ · 22 qDresden segment volumemean |ln(pred/true)| ↓ · 22 qBerlin link volumemean |ln(pred/true)| ↓ · 86 qWhich segment is busieraccuracy ↑ · 63 qDresdenaccuracy ↑ · 54 qMunichMAE ↓ · 66 qMunich motorisation tableaccuracy ↑ · 47 qIntercityBrier ↓ · 42 qKlang ValleyBrier ↓ · 48 qToronto–MontrealBrier ↓ · 45 q
Best observed modelQwen · off0.77Gemma 40.81Qwen · on0.98Qwen · on1.00Qwen · on1.00Qwen · on5.59Qwen · off0.56Llama 3.30.82Qwen · on0.67Qwen · off0.56Llama 3.374.39Qwen · off0.68Gemma 40.68Llama 3.30.15Qwen · on0.56
Naive reference (typical value)not usednot usednot usednot usednot used1.270.570.690.490.5747.85not used0.730.170.61
Logit model (mode choice)not usednot usednot usednot usednot usednot usednot usednot usednot usednot usednot usednot used0.490.110.38
Gemma 4 26B · temperature 00.750.810.910.000.788.160.721.240.190.4481.560.130.680.170.71
Llama 3.3 70B · temperature 00.640.020.550.000.9416.450.940.820.540.4474.390.130.740.150.72
Qwen3.8 27B (thinking off) · temperature 00.770.450.440.120.929.180.561.540.540.56102.030.680.750.260.86
Qwen3.8 27B (thinking on) · temperature 00.530.720.981.001.005.591.151.770.670.33112.150.600.800.220.56
Human referenceNot collected

↑ Higher is better / ↓ Lower is better / “not used”: this reference does not apply to the task / one attempt per question

Details by task type (type scores, references, rounding)
TaskMetricGemma 4Llama 3.3Qwen · offQwen · onReference
Stage 2a: Baseline analysis (graph theory)
Type score · leader: Gemma 4824055740 = naive
Reachability · 371 q · step 2 exact accuracy ↑ 0.75 0.64 0.77 0.53
Maximum flow · 58 q · step 2 exact accuracy ↑ 0.81 0.02 0.45 0.72
Shortest path · 64 q · step 2 exact accuracy ↑ 0.91 0.55 0.44 0.98
Stage 2b: Baseline analysis (local knowledge)
Type score · leader: Qwen · off-31-17-13-180 = naive
Traffic volume · Berlin vehicle-km · 50 q · step 2 accuracy ↑ 0.00 0.00 0.12 1.00
Traffic volume · Dresden station volume · 50 q · step 2 accuracy ↑ 0.78 0.94 0.92 1.00
Traffic volume · Dresden heavy-vehicle share · 22 q · step 2 MAE pp ↓ 8.16 16.45 9.18 5.59 constant 1.27
Traffic volume · Dresden segment volume · 22 q · step 2 mean |ln(pred/true)| ↓ 0.72
×2.1
0.94
×2.6
0.56
×1.8
1.15
×3.1
constant 0.57
Traffic volume · Berlin link volume · 86 q · step 2 mean |ln(pred/true)| ↓ 1.24
×3.5
0.82
×2.3
1.54
×4.7
1.77
×5.9
constant 0.69
Traffic volume · Which segment is busier · 63 q · step 2 accuracy ↑ 0.19 0.54 0.54 0.67 constant 0.49
Station mode · Dresden · 54 q · step 2 accuracy ↑ 0.44 0.44 0.56 0.33 constant 0.57
District motorisation · Munich · 66 q · step 2 MAE ↓ 81.56 74.39 102.03 112.15 constant 47.85
Table information extraction · Munich motorisation table · 47 q · step 2 accuracy ↑ 0.13 0.13 0.68 0.60
Stage 4: Evaluate impact
Type score · leader: Gemma 4-2-4-32-110 = naive
Mode choice prediction · Intercity · 42 q · step 4 Brier ↓ 0.68 0.74 0.75 0.80 constant 0.73
logit 0.49
Mode choice prediction · Klang Valley · 48 q · step 4 Brier ↓ 0.17 0.15 0.26 0.22 constant 0.17
logit 0.11
Mode choice prediction · Toronto–Montreal · 45 q · step 4 Brier ↓ 0.71 0.72 0.86 0.56 constant 0.61
logit 0.38

One attempt per question, temperature 0. All tasks · All questions and answers

Export results

Prompts, answer keys, model responses and scores.