Skip to content

Project record

GovBench

A benchmark of 3,497 multiple-choice questions on UK legislation, parliamentary procedure and government guidance, and the failure that prompted CivBench.

Unverifieddraft
Kind
benchmark (own work)

Recorded because the essay treats it as a productive failure rather than an achievement. Gemma 3 27B scored 94% out of the box, three weeks of fine-tuning gained 1.37 percentage points, and GPT-5 scored 99.26%.

The essay's own verdict is that it measured recall and called it reasoning, and that a model which picks the right option about parliamentary procedure is not a model that can navigate parliamentary procedure. CivBench exists because of that dissatisfaction.

lifecycle: superseded records that relationship in the data rather than only in the prose.

Links to

Referenced by

Source: knowledge/projects/govbench.md

Generated by claude-code/claude-opus-5 on 2026-07-25