Skip to content

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

The collaborative research paper behind CivBench, studying how agents monitor a game and follow through on plans.

Read the source

Article · 2026Sources & confidence: Qualified confidence

Checked 4 October 2026 · AI-assisted review

The collaborative research paper behind CivBench, studying how agents monitor a game and follow through on plans.

What this check doesn’t establish

  • Acceptance is recorded from the author’s announcement; this citation describes the arXiv preprint, not a proceedings version.

Confidence applies to the scope above, not every claim made by the underlying source. How confidence works

The paper introduces a tool-mediated Civilization VI benchmark and studies agents’ monitoring and follow-through. Its 23 admissible runs are a pilot, not a reliable ranking of models.

Written with Austin Tudor David Andrews, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster and Rui Ponte Costa. The work was accepted as a poster at the NeurIPS 2026 Evaluations and Datasets Track. For me, it is a reminder of what becomes possible when you find people who share your curiosity and build together.

The project is the research apparatus; the essay tells the story of the early experiments.

Record history

Catalogue source: knowledge/works/civbench-paper.md

Generated by codex/gpt-6 on 2026-10-04