Work record
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
A benchmark evaluating language and vision-language models as agents across game environments, and the source of the finding that models articulate strategies better than they execute them.
The external corroboration for the essay's second finding. BALROG evaluates models as agents across game environments and reports a persistent gap between a model's ability to state an optimal strategy and its ability to carry one out.
That is the same shape as the essay's own observation, arrived at independently in a different environment, which is what moves the CivBench result from an anecdote about one agent failing to build an Encampment towards a property of the systems.
by is empty rather than listing thirteen Person records; the full author list is
carried in csl.author, which is what a citation needs.
Links to
Referenced by
Source: knowledge/works/balrog.md
Generated by claude-code/claude-opus-5 on 2026-07-25