Settle a question with a research pod
Three members each build and measure one answer at the same time, and the manager writes the decision with a table.
Don't argue about which approach is better. Build all three and measure.
- which rate limiter
- which queue library
- which of three schemas
Design questions usually get settled by whoever argues longest. In this playbook a small team settles one with numbers instead. Each member takes one candidate, builds a small version of it and measures it, all at the same time, and the manager compares the results and writes a decision you can read in a minute and keep in the repository.
Steps
Our repository, gateway, needs per-key rate limiting: 100 requests a minute, short bursts allowed, about 50,000 keys in one process.
- In the repository, run
codeafand send any first message. - Press Alt+V for the wall, press Space on the conversation, then S. In the New team card, clear the name, type
researchand press Enter. - Press Esc, then Alt+2. Press Alt+↑, move to
researchwith ↓, press Enter, then Shift+M. - Send the manager the question:
You manage research for gateway. Question: which rate limiter should gateway.py use? Call team_start three times now, one member per approach, and let them work at the same time: @bucket (token bucket), @slide (sliding window log) and @fixed (fixed window counter). Each builds its approach in research/<its handle>/limiter.py with allow(key), plus a bench.py that measures memory for 50,000 keys and what happens to a burst of 20 requests from one key, runs it, and reports the numbers to you. When all three report, write research/DECISION.md: one table (approach, memory for 50k keys, burst behaviour, lines of code, verdict) and your pick with one paragraph of reasons. Do not change gateway.py.What you see
The header shows ● research ◆ Manager 4 members, and Alt+V followed by the team's number on the Teams row shows its own wall: the manager first, then the three members, each with its numbers. In our run the manager checked on the slowest member once, then wrote the table: the token bucket at about 10 MB for 50,000 keys, the sliding window log at about 42 MB, the fixed window counter at about 9 MB, all three allowing the burst of 20. It picked the token bucket and said why: the fixed window lets up to twice the limit through at a minute boundary, and the sliding log costs four times the memory.
Along the way @bucket found and fixed an off-by-one in its own first version, and the manager reported that too. The prototypes and research/DECISION.md stayed in the repository; gateway.py was untouched.
In our run
About one minute from the question to the decision, for $0.16.
Make it yours
- Any question with candidates. Two queue libraries, three database schemas, four ways to cache a page. One member per candidate.
- Say how to measure. The numbers are only as good as the bench you ask for. Name the load and the limits that matter to you.
- Keep the pod. The team stays in the window. Send the manager the next question, or ask a member to turn the winning prototype into the real change. See Teams.
Go further
Teams: Named, coloured groups of agents, gathered around the outcome they serve.
Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.