FEATUREDComputing Life · Share (鸭哥 research reports)· rssZH00:00 · 06·29
→GraphCast three years later: how Google's weather AI 'breakthrough' was inflated by PR and media
When DeepMind released GraphCast in November 2023, Chinese tech media ran headlines like 'crushed,' 'dominated,' and 'beat the world's best forecast system.' Three years later, ECMWF's public forecast accuracy curve shows zero jump in 2023—it still gains roughly 0.15 days per year. GraphCast's claimed 90% win rate across 1,380 metrics came from counting 6 variables × 37 pressure levels × multiple lead times repeatedly; the hard end metric of 'effective forecast days' didn't budge. None of the 15 Chinese articles quoted an independent meteorologist; in English media, only New Scientist interviewed one, who noted GraphCast doesn't do data assimilation—the most compute-heavy step—and simply feeds on pre-processed data from other systems. Later releases GenCast and WeatherNext 2 repeated the same playbook, with win rates climbing to 97% and 99%, but the baseline quietly shifted from ECMWF's operational system to the team's own previous model. The post suggests three checks when seeing 'crushed' headlines: find the end business metric, count how many quoted sources are truly independent, and verify whether the field has public long-term monitoring data.
#Google DeepMind#GraphCast#ECMWF
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
Three years after GraphCast's '90% win rate' headlines, ECMWF's public forecast accuracy curve shows zero jump in 2023—the metric came from counting 6 variables × 37 pressure levels × multiple lead...
sharp
This piece is worth your time because it reverse-engineers a classic AI PR case using the hardest evidence available: ECMWF's public operational data.
When DeepMind dropped GraphCast in November 2023, Chinese tech media ran with 'crushed' and 'dominated.' But ECMWF's effective forecast days curve—publicly tracked since the 1980s—gains roughly 0.15 days per year, and 2023 shows no jump. Technical Memorandum TM 918 plots GraphCast as a parallel line above the main curve, not a turning point on it.
The '90% win rate across 1,380 metrics'? That's 6 variables × 37 pressure levels × multiple lead times, counted repeatedly. A single-variable advantage gets multiplied across dimensions, inflating the headline number while the end metric—actual forecast days—doesn't budge.
What's more telling: none of the 15 Chinese articles quoted an independent meteorologist. In English media, only New Scientist interviewed one—Ian Renfrew from UEA—who pointed out GraphCast doesn't do data assimilation, the most compute-heavy step (50-67% of the workload). It feeds on pre-processed data from other systems.
Later releases GenCast and WeatherNext 2 pushed win rates to 97% and 99%, but the baseline quietly shifted from ECMWF's operational system to the team's own previous model.
The post leaves you with a practical checklist: when you see 'crushed' headlines, find the end business metric, count truly independent sources, and check whether the field has public long-term monitoring. Weather does. Most AI verticals don't.
HKR breakdown
hook ✓knowledge ✓resonance ✓