Skip to content

Project record

CivBench

An evaluation harness measuring long-horizon strategic competence in language models, using Civilization VI as the environment, with three fixed scenarios and a per-turn agent diary.

Unverifieddraft
Kind
benchmark (own work)
Version
1.0

An evaluation harness built on top of civ6-mcp. Models play full games of up to 330 turns across three fixed scenarios: Ground Control, a fair start; Snowflake, which strands each player on its own arm of a six-armed map; and Cry Havoc.

Each model writes a five-field diary every turn, which serves both as external memory across context compaction and as the record against which stated intentions can later be checked. Without that scaffold only 21% of games reached an ending.

The published pilot covers 23 clean games, weighted heavily towards the easiest scenario, and is presented in the essay as a pilot rather than a ranking.

Recorded as a Project rather than a Work because the numbers it produces are measurements rather than citations. When Claim records are added, the findings drawn from it will point here.

identifiers.swhid is null although the repository holding this code is archived. The Software Heritage snapshot identifies that repository, which is the civ6-mcp record; an identifier for CivBench specifically would need to resolve to a directory within it.

A DOI would be the better identifier here in any case, and none has been minted yet. Until one exists, the repository is what this record should be cited by.

Links to

part of
civ6-mcp

Referenced by

orchestrated by
The Portugal game

Source: knowledge/projects/civbench.md

Generated by claude-code/claude-opus-5 on 2026-07-25