CodeAFThe Model Pool

The Model Pool

Installs measure models on real work; the Model Pool shares the numbers, never the code.

Every finished task adds one honest measurement of which models do the work, shared by everyone and owned by no one.
  • a judge scores each seat
  • numbers leave
  • never code
  • a signed public index

Public benchmarks say how a model does on someone else's tasks. The Model Pool collects how models do on the work people actually hand to CodeAF, seat by seat, and publishes the result as one signed index that every install can read and anyone can check.

Try it in CodeAF

Where

In a terminal.

Type
codeaf pool status

Then press Enter.

You see

This install's pool mode, the cached index and its age, what the last judge did, whether the relay answers, and how many rows wait to be sent, such as pending 0 · can send yes · can read yes.

How a score is made

After a task lands, CodeAF is built to have a model outside the task's crew judge the work, scoring the worker seat and the checker seat when the task had one. codeaf pool status shows whether that has happened yet, as last judge. The judge's call is billed to a judge seat of its own, so its cost shows beside the crew's. The score is kept on this machine, and when the pool is on it also waits in an outbox to be sent.

What leaves your machine

Only text-free numbers leave: the metric, the seat's role, the model, the score, which model judged, the kind of run, a size bucket and the day. Rows are sent under a random per-install id. Code, prompts, file names, paths and anything about you never leave.

A relay adds the rows up and publishes a signed index, which is mirrored on the model-pool branch of the CodeAF repository. codeaf pool verify checks the signature and prints, for example, signature good: version 1790468332, generated 2026-09-27, metrics acceptable, role_quality. The model_pool setting is on by default; read fetches the index and sends nothing, and off turns the pool and its judge off.

Things to build

  • A Monday signature check: a schedule that says "every Monday at 9, run codeaf pool verify and tell me in one line if the signature is not good". Schedules
  • Your own model table: a weekly schedule that reads pool/index.json from the model-pool branch and writes each model's role_quality to MODELS.md in your repo.
  • Read before it leaves: codeaf pool status counts the rows waiting to go, and ~/.codeaf/pool/outbox.jsonl holds the rows themselves: numbers you can read before they are sent.

Go further

Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.