A metric for each type of answer
| Target | Primary metric | Interpretation |
|---|---|---|
| Traffic volume | mean |ln(prediction / reference)| | Lower is better. Exponentiating gives a typical error factor: 1× is exact. Also report MAE in vehicles/day. |
| Heavy-vehicle share | MAE (pp) | Lower is better. Predicting 8% instead of 5% gives an error of 3 percentage points. |
| Choice / platform category | Accuracy + balanced accuracy | Higher is better. Show class frequencies; accuracy alone can favour the majority category. |
| Exclusive mode probabilities | Brier: mean Σ(p − y)² | Lower is better, range 0–2. Probabilities must sum to 1. Secondary: log loss. |
| Independent mode-use probabilities | Brier: mean over people and labels | Lower is better, range 0–1. Probabilities do not need to sum to 1. |
| Policy scenarios | Separate judge rubric | Report each axis and human agreement. No factual-accuracy score or cross-task average. |
These metrics define the new protocol. Historical score formulas and dates remain unchanged. Partial numeric/probability runs receive no primary score; invalid classification answers count as wrong.
Metric definitions: scikit-learn documentation · Research on log accuracy ratios
Baseline results on held-out items
These are computed statistical baselines, not language-model results. Fit uses development items only. Lower loss is better; accuracy is higher-is-better. Brackets show 95% bootstrap intervals.
| Pack | Dev / Test | Primary metric | Constant baseline | Fitted baseline |
|---|---|---|---|---|
| Dresden heavy-vehicle share · extension | 96 / 22 | MAE (percentage points) | 1.273[0.909, 1.636] | Not fitted |
| Dresden stop platforms · transport mode | 186 / 54 | accuracy | 0.574Balanced accuracy: 0.200[0.444, 0.704] | Not fitted |
| Dresden traffic volume · extension | 96 / 22 | mean absolute log ratio | 0.556×1.74[0.391, 0.735] | Not fitted |
| Intercity choice · probability protocol v2 | 168 / 42 | multiclass Brier (sum over classes) | 0.732[0.707, 0.762] | 0.489Regularised conditional multinomial logit[0.390, 0.593] |
| Klang Valley mode use · probability protocol v2 | 152 / 48 | Brier (mean over labels) | 0.178[0.154, 0.203] | 0.121Three independent regularised logistic regressions[0.086, 0.157] |
Held-out sets are small and public. Volume and heavy-vehicle tasks share streets. Mode-choice results describe the selected survey samples, not Dresden or population mode shares.
Existing model answers in raw metrics
Recomputed diagnostics from saved answers; the original result tables are preserved. Intervals reflect variation across questions or streets, not model repeats. Numerical differences alone do not establish a ranking.
Traffic volume · historical 10 segments
| Model | Raw result | 95% interval | Valid / asked |
|---|---|---|---|
| deepseek | 0.712mean absolute log ratio×2.04 | 0.319 – 1.191 | 10 / 10 |
| gemini | 0.346mean absolute log ratio×1.41 | 0.147 – 0.602 | 10 / 10 |
| gemma | 0.735mean absolute log ratio×2.08 | 0.207 – 1.379 | 10 / 10 |
| glm | 0.751mean absolute log ratio×2.12 | 0.304 – 1.342 | 10 / 10 |
| glm53 | 0.711mean absolute log ratio×2.04 | 0.305 – 1.244 | 10 / 10 |
| gpt5 | 0.796mean absolute log ratio×2.22 | 0.336 – 1.357 | 10 / 10 |
| gptoss | 0.837mean absolute log ratio×2.31 | 0.348 – 1.448 | 10 / 10 |
| kimi | 0.781mean absolute log ratio×2.18 | 0.294 – 1.372 | 10 / 10 |
| llama | 0.696mean absolute log ratio×2.00 | 0.278 – 1.193 | 10 / 10 |
| minimax | 0.717mean absolute log ratio×2.05 | 0.329 – 1.257 | 10 / 10 |
| minimax3 | 0.680mean absolute log ratio×1.97 | 0.220 – 1.256 | 10 / 10 |
| qwen | 0.693mean absolute log ratio×2.00 | 0.210 – 1.300 | 10 / 10 |
Heavy-vehicle share · historical 15 segments
| Model | Raw result | 95% interval | Valid / asked |
|---|---|---|---|
| deepseek | 3.767MAE (percentage points) | 1.833 – 6.400 | 15 / 15 |
| gemma | 7.600MAE (percentage points) | 5.833 – 9.500 | 15 / 15 |
| glm | 4.600MAE (percentage points) | 2.267 – 7.200 | 15 / 15 |
| glm53 | 4.333MAE (percentage points) | 2.333 – 7.267 | 15 / 15 |
| gptoss | 6.833MAE (percentage points) | 4.900 – 8.833 | 15 / 15 |
| llama | 7.067MAE (percentage points) | 4.400 – 10.200 | 15 / 15 |
| minimax3 | 4.067MAE (percentage points) | 2.133 – 6.933 | 15 / 15 |
| qwen | 3.600MAE (percentage points) | 1.567 – 6.533 | 15 / 15 |
Traffic choice · historical 12 questions
| Model | Raw result | 95% interval | Valid / asked |
|---|---|---|---|
| gpt5 | 0.583accuracy | 0.250 – 0.833 | 12 / 12 |
| gemini | 0.500accuracy | 0.250 – 0.750 | 12 / 12 |
| kimi | 0.333accuracy | 0.083 – 0.583 | 12 / 12 |
Traffic choice · historical 60 questions
| Model | Raw result | 95% interval | Valid / asked |
|---|---|---|---|
| gemma | 0.417accuracy | 0.300 – 0.533 | 60 / 60 |
| llama | 0.317accuracy | 0.200 – 0.433 | 60 / 60 |
| glm | 0.400accuracy | 0.283 – 0.533 | 60 / 60 |
| minimax | 0.400accuracy | 0.267 – 0.517 | 60 / 60 |
| gptoss | 0.317accuracy | 0.200 – 0.433 | 60 / 60 |