Gemma 4 26B
2of 15 tasks
best model
best model
- 2a Graphs82
- 2b Local knowledge-31
- 4 Mode choice-2
Compare transport knowledge by task. Explore the leading results, then inspect the evidence.
Score per task type: 0 = naive answer, 100 = exact (accuracy × 100 where there is no naive answer).
Every benchmark task belongs to one stage of the planning cycle.
Best result per task is highlighted. Scroll horizontally for all tasks.
| 2a Graphs | 2b Local knowledge | 4 Mode choice | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Traffic volume | Station mode | District motorisation | Table information extraction | Mode choice prediction | |||||||||||
| Model / configuration | Reachabilityexact accuracy ↑ · 371 q | Maximum flowexact accuracy ↑ · 58 q | Shortest pathexact accuracy ↑ · 64 q | Berlin vehicle-kmaccuracy ↑ · 50 q | Dresden station volumeaccuracy ↑ · 50 q | Dresden heavy-vehicle shareMAE pp ↓ · 22 q | Dresden segment volumemean |ln(pred/true)| ↓ · 22 q | Berlin link volumemean |ln(pred/true)| ↓ · 86 q | Which segment is busieraccuracy ↑ · 63 q | Dresdenaccuracy ↑ · 54 q | MunichMAE ↓ · 66 q | Munich motorisation tableaccuracy ↑ · 47 q | IntercityBrier ↓ · 42 q | Klang ValleyBrier ↓ · 48 q | Toronto–MontrealBrier ↓ · 45 q |
| Best observed model | Qwen · off0.77 | Gemma 40.81 | Qwen · on0.98 | Qwen · on1.00 | Qwen · on1.00 | Qwen · on5.59 | Qwen · off0.56 | Llama 3.30.82 | Qwen · on0.67 | Qwen · off0.56 | Llama 3.374.39 | Qwen · off0.68 | Gemma 40.68 | Llama 3.30.15 | Qwen · on0.56 |
| Naive reference (typical value) | not used | not used | not used | not used | not used | 1.27 | 0.57 | 0.69 | 0.49 | 0.57 | 47.85 | not used | 0.73 | 0.17 | 0.61 |
| Logit model (mode choice) | not used | not used | not used | not used | not used | not used | not used | not used | not used | not used | not used | not used | 0.49 | 0.11 | 0.38 |
| Gemma 4 26B · temperature 0 | 0.75 | 0.81 | 0.91 | 0.00 | 0.78 | 8.16 | 0.72 | 1.24 | 0.19 | 0.44 | 81.56 | 0.13 | 0.68 | 0.17 | 0.71 |
| Llama 3.3 70B · temperature 0 | 0.64 | 0.02 | 0.55 | 0.00 | 0.94 | 16.45 | 0.94 | 0.82 | 0.54 | 0.44 | 74.39 | 0.13 | 0.74 | 0.15 | 0.72 |
| Qwen3.8 27B (thinking off) · temperature 0 | 0.77 | 0.45 | 0.44 | 0.12 | 0.92 | 9.18 | 0.56 | 1.54 | 0.54 | 0.56 | 102.03 | 0.68 | 0.75 | 0.26 | 0.86 |
| Qwen3.8 27B (thinking on) · temperature 0 | 0.53 | 0.72 | 0.98 | 1.00 | 1.00 | 5.59 | 1.15 | 1.77 | 0.67 | 0.33 | 112.15 | 0.60 | 0.80 | 0.22 | 0.56 |
| Human reference | Not collected | ||||||||||||||
↑ Higher is better / ↓ Lower is better / “not used”: this reference does not apply to the task / one attempt per question
| Task | Metric | Gemma 4 | Llama 3.3 | Qwen · off | Qwen · on | Reference |
|---|---|---|---|---|---|---|
| Stage 2a: Baseline analysis (graph theory) | ||||||
| Type score · leader: Gemma 4 | 82 | 40 | 55 | 74 | 0 = naive | |
| Reachability · 371 q · step 2 | exact accuracy ↑ | 0.75 | 0.64 | 0.77 | 0.53 | |
| Maximum flow · 58 q · step 2 | exact accuracy ↑ | 0.81 | 0.02 | 0.45 | 0.72 | |
| Shortest path · 64 q · step 2 | exact accuracy ↑ | 0.91 | 0.55 | 0.44 | 0.98 | |
| Stage 2b: Baseline analysis (local knowledge) | ||||||
| Type score · leader: Qwen · off | -31 | -17 | -13 | -18 | 0 = naive | |
| Traffic volume · Berlin vehicle-km · 50 q · step 2 | accuracy ↑ | 0.00 | 0.00 | 0.12 | 1.00 | |
| Traffic volume · Dresden station volume · 50 q · step 2 | accuracy ↑ | 0.78 | 0.94 | 0.92 | 1.00 | |
| Traffic volume · Dresden heavy-vehicle share · 22 q · step 2 | MAE pp ↓ | 8.16 | 16.45 | 9.18 | 5.59 | constant 1.27 |
| Traffic volume · Dresden segment volume · 22 q · step 2 | mean |ln(pred/true)| ↓ |
0.72 ×2.1 |
0.94 ×2.6 |
0.56 ×1.8 |
1.15 ×3.1 |
constant 0.57 |
| Traffic volume · Berlin link volume · 86 q · step 2 | mean |ln(pred/true)| ↓ |
1.24 ×3.5 |
0.82 ×2.3 |
1.54 ×4.7 |
1.77 ×5.9 |
constant 0.69 |
| Traffic volume · Which segment is busier · 63 q · step 2 | accuracy ↑ | 0.19 | 0.54 | 0.54 | 0.67 | constant 0.49 |
| Station mode · Dresden · 54 q · step 2 | accuracy ↑ | 0.44 | 0.44 | 0.56 | 0.33 | constant 0.57 |
| District motorisation · Munich · 66 q · step 2 | MAE ↓ | 81.56 | 74.39 | 102.03 | 112.15 | constant 47.85 |
| Table information extraction · Munich motorisation table · 47 q · step 2 | accuracy ↑ | 0.13 | 0.13 | 0.68 | 0.60 | |
| Stage 4: Evaluate impact | ||||||
| Type score · leader: Gemma 4 | -2 | -4 | -32 | -11 | 0 = naive | |
| Mode choice prediction · Intercity · 42 q · step 4 | Brier ↓ | 0.68 | 0.74 | 0.75 | 0.80 | constant 0.73 logit 0.49 |
| Mode choice prediction · Klang Valley · 48 q · step 4 | Brier ↓ | 0.17 | 0.15 | 0.26 | 0.22 | constant 0.17 logit 0.11 |
| Mode choice prediction · Toronto–Montreal · 45 q · step 4 | Brier ↓ | 0.71 | 0.72 | 0.86 | 0.56 | constant 0.61 logit 0.38 |
One attempt per question, temperature 0. All tasks · All questions and answers
Prompts, answer keys, model responses and scores.