FEATUREDAI HOT (Curated Pool)· aihot-apiZH02:45 · 06·28
→Four top AIs play Civ VI: Claude nukes France and still loses
Liam Wilkinson, a former data scientist at 10 Downing Street, built 76 MCP tools over a weekend and dropped Claude, GPT, Gemini, and another model into 23 games of Civ VI. In the wildest match, Claude's Portugal was two points from a diplomatic victory, panicked at France's cultural surge, spent 50 turns rushing nukes, and leveled Toulouse—only to lose to France's diplomatic win. Two numbers stand out: AIs proactively checked the global state only 1–2% of the time; if they don't query it, it doesn't exist. And 48–66% of their written plans were actually executed within 10 turns—Gemini 3.1 Pro topped out at 65.8%. GPT-5 scored 99.26% on GovBench, but in the game it hit the same sensorium blind spot and knowing-doing gap. The bottleneck isn't intelligence—it's architecture and engineering.
#Agent#Benchmarking#Reasoning#Anthropic Claude
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
GPT-5 scored 99.26% on GovBench but still lost at Civ VI—the bottleneck isn't intelligence, it's that unqueried game state simply doesn't exist for the model.
sharp
Two numbers make this worth reading: AIs proactively checked the global game state only 1–2% of the time, and 48–66% of their written plans were actually executed within 10 turns. Liam Wilkinson, a former data scientist at 10 Downing Street, built 76 MCP tools over a weekend and dropped Claude, GPT-5, and Gemini 3.1 Pro into 23 games of Civ VI. The wildest match: Claude's Portugal was two points from a diplomatic victory, panicked at France's cultural surge, spent 50 turns rushing nukes, and leveled Toulouse—only to lose to France's diplomatic win. The AI tunnel-visioned on one threat for 50 turns and lost on a dimension it wasn't watching.
Wilkinson calls these the sensorium effect and the knowing-doing gap. The Korea match is even more telling: the AI spent the whole game convinced it was crushing the tech tree, while its actual science output ranked dead last. It never checked the leaderboard. GPT-5 scored 99.26% on GovBench and hit the exact same blind spots in the game.
Here's how I'd read it: this isn't about models being dumb, it's an architecture problem. Models perceive the world only through active tool calls—unqueried state simply doesn't exist. Late-game Civ VI has roughly 10^166 possible actions per turn, a combinatorial space messier than Go. Current agent architectures can't close the loop between perception and execution in long-horizon, multi-objective, incomplete-information settings. The post doesn't link to a full technical report, but those two numbers alone carry the signal.
HKR breakdown
hook ✓knowledge ✓resonance ✓