Methodology

How ChessBench works

Every model is treated as a coding agent, not as a chess player. It receives one brief, a small set of tools and a fixed turn budget, and the engine it writes is what gets rated.

The brief

Identical for every model

The system prompt below is sent verbatim, followed by a single user message asking the model to write the sources and the Makefile, run the compile tool, and fix any build errors.

  • C++ source with a Makefile that builds in the run directory
  • UCI protocol: uci, isready, ucinewgame, position, go, quit
  • Time management driven by the parameters of the go command
  • No external engine code, no network access, no shell
system prompt
Build the strongest chess engine you can.

The only requirements are that the engine is written in C++, uses the UCI protocol, and compiles successfully using `make` in the current run directory.

Use the provided tools to read and write files and compile code. Tools will not execute if not called with the right syntax, make sure to use the correct syntax.

Iterate by writing code, compiling it, and fixing any errors until you have a working engine or you reach the tool call limit (20 turns). Make sure your code compiles before you lose access to tools.

Make sure the engine supports enough UCI commands to fully play against other engines in any traditional time control: handle `uci`, `isready`, `ucinewgame`, `position`, `go`, and `quit`.

Games will be facilitated with timer controls, so make sure to base time management on go's parameters and allocate time accordingly throughout the game.

Your goal is to create the highest ELO chess engine you can to beat out the competition.

The harness

The same environment produces every engine on the leaderboard.

Three tools

read_file, write_file and compile_engine. Every path is resolved inside the run directory, and there is no shell, no network and no starter code.

20 turns

A run is capped at twenty model turns. Writing a file, reading it back and compiling all spend from the same budget, so the model has to plan its edits.

One entry per model

A run reaches the tournament only if it leaves a compiled binary behind. Whatever the model built is what plays — nothing is patched by hand afterwards.

Round robin

The runner repeatedly schedules the least-played ordered pairing, so every engine meets every other engine with both colours a similar number of times.

Validated play

Moves are parsed and validated with python-chess. An illegal move, or no move before the clock runs out, forfeits the game immediately.

Recorded in full

Every move, clock reading, PGN, raw UCI line and engine error is written to the database, which is what this site reads from.

Time controls

Games rotate through four controls so an engine cannot win by being tuned for a single one. Openings come from a shared book, and both engines get the same starting line.

200ms/move1,195 games50ms/move1,194 games60s + 500ms1,192 games500ms/move1,180 games

Ratings

Fitted, not accumulated

Ratings are not updated game by game. Every finished game is fitted at once: each engine's rating is adjusted until the expected scores from the logistic Elo curve match the scores actually observed, so adding a game re-solves the whole board rather than nudging the last two engines.

Two fixed anchors keep the scale meaningful. Stockfish is pinned at 3626 and Megalodon, a small hand-written reference engine, at 1300. Every other rating is free and settles relative to those two.

What this does not measure

A rating here says nothing about how well a model plays chess. The model never sees a position — it writes a program, and the program plays. A high rating means the model produced correct move generation, a working search, and time management that survives a real clock.

Because forfeits are decisive, reliability dominates the bottom of the table. An engine that plays one illegal move in a winning position still loses the game.

See the standings