Fire season active. Emergency resources & evacuation info →
FEATURED: Human emissions have likely delayed the next glacial period by tens of thousands of years. Read the full story.

If the Model Fails, You’ll See It Failing Here

Updated June 24, 2026 · active

This page is the platform’s accountability mechanism: a public, pre-registered validation of the Firewatch model during the 2026 fire season. The rules for scoring are written down below, before the outcomes are known. Predictions are timestamped when they are issued, and they get scored against what actually burned, whether that makes the model look good or not. A model that is only graded by the people who built it, after the fact, hasn’t been graded at all.

The Scoring ProtocolPre-registered · 2026 season

Once daily forecasts are live, each day the model issues corridor-wide rankings of every 4km cell. Each prediction is timestamped before outcomes are known. Monthly during the season, and again at season end, we score three things:

Hit rate vs. climatology. What fraction of new large fires (>1,000 acres) ignited in the top 10% and top 25% of ranked cells, compared against a climatological baseline built from fire history 2000–2024. Beating the baseline is the bar; matching it means the model adds nothing.
Calibration of published probabilities. If we publish a probability, it has to mean something. A reliability (calibration) curve checks whether cells given, say, a 10% chance actually burn about 10% of the time. Overconfident probabilities get reported as overconfident.
Misses, listed by name. Every large fire that ignited outside the model's top-ranked cells gets named here, with where the model ranked that cell. No aggregating misses away into a summary statistic.
Scored PredictionsUpdated June 24, 2026 · active
DatePredictionOutcomeScore
2026-06-24Firewatch v1.3 daily snapshot
2026-06-23Firewatch v1.3 daily snapshot
2026-06-22Firewatch v1.3 daily snapshot
2026-06-18Firewatch v1.3 daily snapshot
2026-06-17Firewatch v1.3 daily snapshot
2026-06-16Firewatch v1.3 daily snapshot
2026-06-15Firewatch v1.3 daily snapshot
2026-06-14Firewatch v1.3 daily snapshot
2026-06-13Firewatch v1.3 daily snapshot
What We Got Wrong So FarA changelog of corrections

An accountability page that only starts counting once the model is fixed isn’t accountability. These are the mistakes we have already made, what they affected, and what we did about them. The list will grow.

2025Corrected

v0.7 published a leaked metric

The v0.7 AUC of 0.913 was computed on a test set that had been used during model tuning, which inflates the number. We disclosed it on the Methods model-history table (it carries an asterisk there), rebuilt the evaluation with a proper three-way train/validation/test split in v0.8, and reported the lower, honest number.

Early 2026Corrected

“91% accuracy” misphrasing in early posts

Early posts described the model's 0.911 AUC-ROC as “91% accuracy.” That is wrong: AUC-ROC is a ranking metric, not a hit rate. Given one cell that burned and one that didn't, the model ranks the burned cell higher 91% of the time, at a 0.1% base rate, most flagged cells will not burn. The phrasing has been corrected and Methods now explains the metric in plain language.

June 2–11, 2026Corrected

Nine-day data outage during active fire season

Daily updates stopped on June 2 and did not resume until June 11, with fire season underway, the site served stale active-fire and conditions data for nine days without a visible warning. Updates have resumed. The deeper failure was not flagging staleness to readers; surfacing data age prominently is now on the fix list.

June 2026Open

Headline 0.911 AUC suspected to be inflated

An internal audit found that the v2.1 evaluation allowed spatial memorization (the model could learn which cells burn) and several forms of feature leakage. The v2.2 re-validation uses spatially blocked held-out testing and fixes the leaks; the honest number is expected to be lower. Until it lands, treat 0.911 as an upper bound, not a result. The corrected figure will be published here and on Methods, whatever it is.

Scoring data lives in a public file (/scorecard.json) so anyone can check our arithmetic. Model details, known limitations, and the full version history are on Methods. Live fire weather and the seasonal risk surface are on the Forecast page. Found an error we haven’t listed? Tell us. it goes on the list.