AI4Fire: no LLM beat repeating today’s wildfire staffing count
A new wildfire benchmark ran 35 language models over five fire tasks. Its one headline gain measures database access, not what the models know. Give a model a read-only SQL tool and its accuracy on questions about US fire records jumps from at most 0.160 to 0.885 or better. On the staffing task, by contrast, no core model beat a rule as simple as repeating today’s personnel count.
AI4Fire went up on arXiv on 7 October. The five authors are led by Yue Zhao of the University of Southern California, with co-authors at Arizona State, UNLV and the University of Maine. Six core models run every item twice, bare and grounded, across 1,446 paired items, and a sweep adds 29 more systems. That makes 35 models from twelve vendors, and the team released all 77,644 scored responses with a script that rebuilds the paper’s tables offline.
The core six are claude-opus-5, claude-opus-4.8, gemini-3.1-pro, gpt-6-astra, Qwen3-VL-235B-A22B and Llama 4 Maverick. Grounded means one task-specific addition, and only two of the five additions can carry an item’s answer. The database task asks 156 questions over 2,303,566 wildfire occurrence records. Its grounded arm may fire up to eight read-only queries at a frozen snapshot, which is where the gains of 0.808 to 0.994 come from.
| Task and metric | Non-LLM comparator | Core six, bare | Core six, grounded |
|---|---|---|---|
| Fire record questions, accuracy | Best constant 0.141 | 0.000 to 0.160 | 0.885 to 1.000 |
| Next-day staffing, normalised error (lower is better) | Persistence 0.146 | 0.157 to 0.193 | 0.170 to 0.248 |
| Aerial question answering, accuracy | Majority class 0.627 | 0.534 to 0.650 | 0.544 to 0.657 |
| Fire danger forecasting, AUPRC | Temperature rule 0.654 | 0.485 to 0.697 | 0.515 to 0.682 |
| Smoke detection, recall | Frame difference 0.920 | 0.509 to 0.643 | 0.554 to 0.741 |
Even there the open-weight pair trails, at 0.885 and 0.891 against 0.994 to 1.000 for the proprietary four. Of their 35 remaining failures, 30 were dates filtered in a guessed format. So the task separates models on plumbing rather than on fire knowledge.
The staffing task is the clearest loss. Persistence, which just repeats today’s filed personnel count, scores 0.146 normalised error, but no core model got under it. A small trained regressor matched the rule at 0.144.
Handing the models six similar earlier fire-days made things worse on five of six. That’s because the two open-weight models copied the median of those analogues on 284 and 259 of 300 items. The copying rule forecasts at 0.249 on its own, so grounding there bought a worse answer.
Zero-shot prompted models rarely beat a trivial rule or a trained comparator, and every task now carries one.
Zhao et al., AI4Fire
Three public fire datasets carry defects
Rebuilding the tasks turned up a second finding, which is that some public releases score the dataset rather than the model. In the Mesogeos Track A fire-danger files, a static column called burned_area_has is nonzero for 100 percent of holdout positives and zero percent of negatives. That alone separates the holdout at average precision 1.000.
Two further columns leak too, though less cleanly. The column burned_areas is nonzero on 67.06 percent of test positives and no negatives, while ignition_points is nonzero on 26.15 percent.
Track A evaluations have to withhold all three columns and supply only the 24 runnable driver features, the paper says. Mesogeos documents Track A as binary classification, and it lists rasterised burned areas and ignition points among its datacube variables. So the columns are no secret. What the audit adds is that they survive into the released split.
The aerial task has a different problem. On 67 of 738 Sycan Marsh frames the thermal maximum is exactly 187.424 degrees Celsius. The released labels read that clipped value as smoldering or fire-free, though FLAME 3 files all 67 frames under Fire. Fifteen of the core six’s 34 grounded misses on closed-form items sit on those frames.
A third source dropped out of the benchmark entirely. The team withdrew Fire360 because none of its 100 released prompts names a video file, frame number or timestamp.
Calibration matters more than the ranking
Fire agencies already ship this kind of tool. California launched Ask CAL FIRE, a chatbot answering wildfire questions in 70 languages, in May 2025. On the fire-danger sample, all twelve core-model runs stated probabilities less calibrated than a calendar-month prior. Their expected calibration errors ran 0.119 to 0.326 against 0.037 for that prior.
Two core models called fire on 4 and 6 percent of items when the sample’s own positive rate was 34 percent, and the prompt never stated that rate. The gap between a leaderboard place and a usable forecast is the pattern this desk keeps finding, from ORCA-bench, where GLM-5 invented a root cause in 66 percent of hard incidents, to Reflection AI’s Beam trailing two rivals on its own benchmark.
Next up, the authors say, is multi-step agent tasks over the same public fire databases, scored beside trivial and trained comparators. Until then, anyone citing a wildfire model score has a cheap check available. Both the comparator and the raw responses sit in the repository.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
