Essay record
I Gave an AI a Civilization to Run. It Built a Nuke.
An account of building an evaluation harness around Civilization VI, and of two failures that showed up across every model tested: not looking at what it had not thought to ask for, and not doing what it had written down.
This is a catalogue record describing a published essay, listing what the essay draws on and what has been checked. It is not the essay.
Read the essay →published 2026-06-22
The post that gives the collection its clearest example of measurement as distinct from citation. Its findings were produced by apparatus its author built, and it reports them as a pilot rather than a ranking.
The argument runs from a benchmark that failed by succeeding, GovBench, through a keyhole into a game engine, to two failures that held across every model: the sensorium effect, where an agent goes blind to whatever it does not think to ask about, and the knowing-doing gap, where it writes down the right move and does not make it.
Findings this essay reports
These are measurements produced by CivBench over 23 clean games, not claims drawn from sources. The collection currently has no type for them, which is discussed below.
- Plan follow-through within ten turns: Claude Opus 4.6 48.2%, GPT-5.4 63.2%, Gemini 3.1 Pro 65.8%
- In 7 of 20 losses where a rival's victory was visible in advance, the agent never checked for it in the twenty turns before losing
- Whole-board checks account for 1 to 2% of agent actions; against an instruction to check every twenty turns, models managed four to ten checks across a 330-turn game rather than about sixteen
- Without the external diary, only 21% of games reached an ending
The essay states its own limits plainly: 23 games is a pilot, and nearly all of them are the gentlest of the three scenarios.
Findings are now records
This essay was the trigger case for the Finding type, which was promoted on
25 July 2026 ahead of Claims rather than alongside them. Each number above is now a
record carrying its method, its sample, the apparatus that produced it and the
caveats it comes with, so the collection can ask which findings rest on 23 games
rather than only display them.
They are listed in reports. The prose above is kept because a reader should not
have to traverse seven records to learn what the essay found.
What this record does not capture
discusses is empty pending Concept records. This essay would contribute several
of its own, and the sensorium effect is the clearest coinage in the five posts.
examines names the Portugal game, which
is the first case recorded under the agent-run payload. Until that payload
existed the game could not be filed at all: it has actors, a plan, a decision under
constraint and an outcome, but no jurisdiction, no policy domain and no wave of
technocratic thought, and the only payload in use required all three.
The other games in the pilot are not recorded individually. The aggregate results are already carried by the findings, and a case record per game would restate those numbers without adding anything. The Portugal game earns one because the essay reads it closely enough to be checkable.
Links to
- draws on
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesModel Context ProtocolDr. Strangelove or: How I Learned to Stop Worrying and Love the Bomb
- examines
- The Portugal game
- illustrated by
- The nuking of ToulouseWhere the agent spent its attention, Korea on SnowflakeThe Quarter Quell arenaNot an exercise, Dr. StrangeloveMajor Kong rides the bomb, Dr. StrangeloveThis Is FineMachiavelli, reputation versus realityGromit laying track ahead of the trainWallace in the wrong trousers, landing on the trainLiftoff, A Grand Day Out
- mentions
- Liam WilkinsonStanley Kubrick
- reports
- Plan follow-through, Claude Opus 4.6Plan follow-through, GPT-5.4Plan follow-through, Gemini 3.1 ProLosses where an imminent rival victory was never checkedShare of agent actions spent checking the whole boardVictory-condition checks performed against the sixteen instructedGames reaching an ending without the external diary
Referenced by
- cited in
- civ6-mcpCivBenchSid Meier's Civilization VIGovBenchBALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesDr. Strangelove or: How I Learned to Stop Worrying and Love the BombModel Context Protocol
- examined in
- The Portugal game
- reported in
- Games reaching an ending without the external diaryPlan follow-through, Claude Opus 4.6Plan follow-through, Gemini 3.1 ProPlan follow-through, GPT-5.4Losses where an imminent rival victory was never checkedVictory-condition checks performed against the sixteen instructedShare of agent actions spent checking the whole board
- used in
- The nuking of ToulouseLiftoff, A Grand Day OutGromit laying track ahead of the trainWhere the agent spent its attention, Korea on SnowflakeMachiavelli, reputation versus realityThe Quarter Quell arenaNot an exercise, Dr. StrangeloveMajor Kong rides the bomb, Dr. StrangeloveThis Is FineWallace in the wrong trousers, landing on the train
Source: knowledge/essays/i-gave-an-ai-a-civilization.md
Generated by claude-code/claude-opus-5 on 2026-07-25