The Model Pool
Installs measure models on real work; the Model Pool shares the numbers, never the code.
Every finished task adds one honest measurement of which models do the work, shared by everyone and owned by no one.
- a judge scores each seat
- numbers leave
- never code
- a signed public index
Public benchmarks say how a model does on someone else's tasks. The Model Pool collects how models do on the work people actually hand to CodeAF, seat by seat, and publishes the result as one signed index that every install can read and anyone can check.
Try it in CodeAF
- Where
In a terminal.
- Type
codeaf pool statusThen press Enter.
- You see
This install's pool mode, the cached index and its age, what the last judge did, whether the relay answers, and how many rows wait to be sent, such as
pending 0 · can send yes · can read yes.
How a score is made
After a task lands, CodeAF is built to have a model outside the task's crew judge the work, scoring the worker seat and the checker seat when the task had one. codeaf pool status shows whether that has happened yet, as last judge. The judge's call is billed to a judge seat of its own, so its cost shows beside the crew's. The score is kept on this machine, and when the pool is on it also waits in an outbox to be sent.
What leaves your machine
Only text-free numbers leave: the metric, the seat's role, the model, the score, which model judged, the kind of run, a size bucket and the day. Rows are sent under a random per-install id. Code, prompts, file names, paths and anything about you never leave.
A relay adds the rows up and publishes a signed index, which is mirrored on the model-pool branch of the CodeAF repository. codeaf pool verify checks the signature and prints, for example, signature good: version 1790468332, generated 2026-09-27, metrics acceptable, role_quality. The model_pool setting is on by default; read fetches the index and sends nothing, and off turns the pool and its judge off.
Things to build
- A Monday signature check: a schedule that says "every Monday at 9, run codeaf pool verify and tell me in one line if the signature is not good". Schedules
- Your own model table: a weekly schedule that reads
pool/index.jsonfrom themodel-poolbranch and writes each model'srole_qualityto MODELS.md in your repo. - Read before it leaves:
codeaf pool statuscounts the rows waiting to go, and~/.codeaf/pool/outbox.jsonlholds the rows themselves: numbers you can read before they are sent.
Go further
Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.