Skip to content
← All work
2026ShippedPublic

Biome

An evaluation lab I built as a strategy game: I put local and cloud models on a live ELO ladder, let a simulation judge every move, and trained the next generation on the best ones.

Solo — game design, simulation, AI, build

Models rated on the ladder
44
Rated matches recorded
378
Engine-graded turns
15,168
Species across three trophic tiers
6

01 / problem

The problem

Benchmarks tell you how a model does on a test set. I wanted to see how local models behave when they have to make a sequence of spatial, strategic decisions and live with the consequences, and to compare them fairly over hundreds of games instead of a handful of vibes.

That needed a game where the outcome is decided by a simulation rather than by another model's opinion, so every result is objective and repeatable.

02 / build

What I built

A browser game backed by a small Python server and Ollama, with no build step. Local weights and Ollama's cloud-routed models play on the same ladder, and AWS Bedrock is an optional third provider.

  • An ecosystem engine. Six species across plants, herbivores and predators on a procedurally generated hex map. Energy enters only through the soil and returns through decomposition, and the ideal trophic ratios are derived from the energy economy rather than hand-tuned.
  • Swappable board vision. Four map strategies (region digest, ASCII, layered ASCII, raw coordinates) share one registry, so a match isolates the effect of how the board is described. I measured the prompt cost of each, and fog of war holds in every one.
  • Robust model play. A three-tier JSON extractor handles chatty replies, and a deterministic fallback means a bad answer never stalls a match.
  • A persistent ladder. The server owns the ELO math (K=32, base 1000) in SQLite, with 8/16/32-model tournaments in six draw formats and per-model dossiers.
  • A broadcast, not a log. Watch mode is a live spectator view: fighter dossiers, odds, event banners (extinctions, trophic chains), a score timeline, and the two models trading commentary about each other's play. Every model has a ComfyUI-generated persona and animations.
  • A training pipeline. Every turn is a clean prompt-to-move pair, the engine grades each move, and the best moves become a dataset for LoRA-distilling a champion's play into a smaller model, exported to GGUF and put back on the ladder.

03 / signal

What it shows

Systems design for LLM evaluation: an objective judge, a controlled variable (board presentation), and a ladder as the eval metric. It also shows care for fairness details that are easy to miss, like shuffled simulation order and fog of war enforced in the prompt itself, not just in the UI.

And it is a complete loop, from generating data through self-play to training, exporting and re-evaluating a fine-tuned model on the same ladder. The loop works: a distilled 14B of my own sits fourth of 44 on the ladder, ahead of most of the cloud models it learned alongside.

Gallery 1 / 7

Match, ladder, training and vision lab in one pass (sped up, no sound).