Skip to content

Work record

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

A benchmark evaluating language and vision-language models as agents across game environments, and the source of the finding that models articulate strategies better than they execute them.

Unverifieddraft

The external corroboration for the essay's second finding. BALROG evaluates models as agents across game environments and reports a persistent gap between a model's ability to state an optimal strategy and its ability to carry one out.

That is the same shape as the essay's own observation, arrived at independently in a different environment, which is what moves the CivBench result from an anecdote about one agent failing to build an Encampment towards a property of the systems.

by is empty rather than listing thirteen Person records; the full author list is carried in csl.author, which is what a citation needs.

Links to

Referenced by

Source: knowledge/works/balrog.md

Generated by claude-code/claude-opus-5 on 2026-07-25