Pareto crewing
Pick a model for each seat, per task, on the cost-quality frontier.
Route the crew, not the call. A task is a team of seats, and each seat deserves its own model.
- cheap crew for a bugfix
- stronger checker for open-ended work
- --best for one task
A task in CodeAF runs on a crew of three seats: the worker that does the work, the planner that structures it, and the checker that reads the result before it lands. Each seat can run on a different model, and by default all three are on auto, so CodeAF picks them again for every task.
Try it in CodeAF
- Where
In any conversation.
- Type
/crewThen press Enter.
- You see
The three seats on auto, each with the model it is likely to get, such as
auto · likely glm-5.3.
How the crew is picked
CodeAF first reads what kind of work the task is: a bugfix, a complex fix, open-ended work, or other. Then it scores every allowed model for each seat on that kind of work, weighs the score against the price, and stops at the knee, where more money stops buying much. A small fix usually runs on a cheap crew. On open-ended work the checker, the seat that accepts the work, is upgraded first. When a task starts, one line names the crew it got, for example task 4 crew · other · worker glm-5.3 (openrouter) · checker claude-opus-5.5, then its cost.
Each task is held to its own limit, $5 unless you change it in /crew. To buy more for one task only, start it with /task --best and every seat is upgraded for that task.
Why seats, not calls
Routing one model per call puts the same model in every seat. The best crews in our study were mixed. On real GitHub issues, a class-aware crew matched the best fixed crew's quality at 40% lower cost. On open-ended work, upgrading a few seats of a cheap crew took mergeable fixes from 0 of 8 to 8 of 8.
Things to build
- Small fixes under 50 cents:
/crew cap task 0.5, then hand off small fixes: a crew estimated over the cap is swapped for one that fits, and the crew line says so. - A trial week for a new worker:
/crew pin worker z-ai/glm-5.3-flashputs it in the worker seat for every task while the planner and checker stay on auto;/crew unpin workerends the trial. Choose the model for a task - Cheap first, stronger on demand: let auto pick the crew, and when a result falls short,
/redo strongerruns the same task again with the one seat that buys the most moved up a rung. - A model bake-off: run the same brief twice, "as a task on kimi k3" and "as a task on glm 5.3 flash", and compare the two task pages' cost and result.
Go further
Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.