CodeAFSettle a question with a research pod

Settle a question with a research pod

Three members each build and measure one answer at the same time, and the manager writes the decision with a table.

Don't argue about which approach is better. Build all three and measure.
  • which rate limiter
  • which queue library
  • which of three schemas

Design questions usually get settled by whoever argues longest. In this playbook a small team settles one with numbers instead. Each member takes one candidate, builds a small version of it and measures it, all at the same time, and the manager compares the results and writes a decision you can read in a minute and keep in the repository.

Steps

Our repository, gateway, needs per-key rate limiting: 100 requests a minute, short bursts allowed, about 50,000 keys in one process.

  1. In the repository, run codeaf and send any first message.
  2. Press Alt+V for the wall, press Space on the conversation, then S. In the New team card, clear the name, type research and press Enter.
  3. Press Esc, then Alt+2. Press Alt+↑, move to research with ↓, press Enter, then Shift+M.
  4. Send the manager the question:
You manage research for gateway. Question: which rate limiter should gateway.py use? Call team_start three times now, one member per approach, and let them work at the same time: @bucket (token bucket), @slide (sliding window log) and @fixed (fixed window counter). Each builds its approach in research/<its handle>/limiter.py with allow(key), plus a bench.py that measures memory for 50,000 keys and what happens to a burst of 20 requests from one key, runs it, and reports the numbers to you. When all three report, write research/DECISION.md: one table (approach, memory for 50k keys, burst behaviour, lines of code, verdict) and your pick with one paragraph of reasons. Do not change gateway.py.

What you see

The header shows ● research ◆ Manager 4 members, and Alt+V followed by the team's number on the Teams row shows its own wall: the manager first, then the three members, each with its numbers. In our run the manager checked on the slowest member once, then wrote the table: the token bucket at about 10 MB for 50,000 keys, the sliding window log at about 42 MB, the fixed window counter at about 9 MB, all three allowing the burst of 20. It picked the token bucket and said why: the fixed window lets up to twice the limit through at a minute boundary, and the sliding log costs four times the memory.

Along the way @bucket found and fixed an off-by-one in its own first version, and the manager reported that too. The prototypes and research/DECISION.md stayed in the repository; gateway.py was untouched.

In our run

About one minute from the question to the decision, for $0.16.

Make it yours

  • Any question with candidates. Two queue libraries, three database schemas, four ways to cache a page. One member per candidate.
  • Say how to measure. The numbers are only as good as the bench you ask for. Name the load and the limits that matter to you.
  • Keep the pod. The team stays in the window. Send the manager the next question, or ask a member to turn the winning prototype into the real change. See Teams.

Go further

Teams: Named, coloured groups of agents, gathered around the outcome they serve.

Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.