Three tools
read_file, write_file and compile_engine. Every path is resolved inside the run directory, and there is no shell, no network and no starter code.
Methodology
Every model is treated as a coding agent, not as a chess player. It receives one brief, a small set of tools and a fixed turn budget, and the engine it writes is what gets rated.
The brief
The system prompt below is sent verbatim, followed by a single user message asking the model to write the sources and the Makefile, run the compile tool, and fix any build errors.
Build the strongest chess engine you can. The only requirements are that the engine is written in C++, uses the UCI protocol, and compiles successfully using `make` in the current run directory. Use the provided tools to read and write files and compile code. Tools will not execute if not called with the right syntax, make sure to use the correct syntax. Iterate by writing code, compiling it, and fixing any errors until you have a working engine or you reach the tool call limit (20 turns). Make sure your code compiles before you lose access to tools. Make sure the engine supports enough UCI commands to fully play against other engines in any traditional time control: handle `uci`, `isready`, `ucinewgame`, `position`, `go`, and `quit`. Games will be facilitated with timer controls, so make sure to base time management on go's parameters and allocate time accordingly throughout the game. Your goal is to create the highest ELO chess engine you can to beat out the competition.
The same environment produces every engine on the leaderboard.
read_file, write_file and compile_engine. Every path is resolved inside the run directory, and there is no shell, no network and no starter code.
A run is capped at twenty model turns. Writing a file, reading it back and compiling all spend from the same budget, so the model has to plan its edits.
A run reaches the tournament only if it leaves a compiled binary behind. Whatever the model built is what plays — nothing is patched by hand afterwards.
The runner repeatedly schedules the least-played ordered pairing, so every engine meets every other engine with both colours a similar number of times.
Moves are parsed and validated with python-chess. An illegal move, or no move before the clock runs out, forfeits the game immediately.
Every move, clock reading, PGN, raw UCI line and engine error is written to the database, which is what this site reads from.
Games rotate through four controls so an engine cannot win by being tuned for a single one. Openings come from a shared book, and both engines get the same starting line.
Ratings are not updated game by game. Every finished game is fitted at once: each engine's rating is adjusted until the expected scores from the logistic Elo curve match the scores actually observed, so adding a game re-solves the whole board rather than nudging the last two engines.
Two fixed anchors keep the scale meaningful. Stockfish is pinned at 3626 and Megalodon, a small hand-written reference engine, at 1300. Every other rating is free and settles relative to those two.
A rating here says nothing about how well a model plays chess. The model never sees a position — it writes a program, and the program plays. A high rating means the model produced correct move generation, a working search, and time management that survives a real clock.
Because forfeits are decisive, reliability dominates the bottom of the table. An engine that plays one illegal move in a winning position still loses the game.
See the standings