Headless mode
Headless mode runs one job with nobody watching, then exits with an answer and an exit code.
The same agent, with the conversation taken out.
- one brief in, one JSON object out
- exit 0 done, 2 did not finish, 3 a limit stopped it
- nothing stops to ask
Try it in CodeAF
- Where
In a terminal, inside any git repo.
- Type
codeaf do "Make the failing tests pass by fixing the code, not the tests." --json- You see
One JSON object on stdout with
ok,stop,answer,filesandspend_usd, and$?is 0.
Built for nobody watching
codeaf do "<brief>" is not the chat with the window hidden. The chat is made for talking: it carries the conversation, names the chat, routes each turn, and can stop to ask you. Headless drops all of that. The brief goes straight to a planner, a worker and a checker, and the worker's prompt has no ask tool, no approval cards and no history. In our run of the same small fix, the chat's first call sent about 14,700 prompt tokens with 26 tools; do sent about 6,200 with 8.
Three verbs, one contract
do plans, may split the job, and has a checker judge the result. exec runs one worker for one pass with no plan and no checker. run follows a program somebody saved. All three edit the folder in place, print the answer on stdout and everything else on stderr, and exit the same way: 0 done, 1 could not run, 2 did not finish, 3 a limit you set stopped it. A do that hits --timeout exits 124.
It decides alone, inside your limits
None of them takes --yolo, because nothing in them stops to ask. What do still refuses is spend: every call is priced first, and a plan that would cross your daily limit stops unless you passed --yes-spend. Choose the crew for one run with --cheap, --best or --pin worker=<model>, and cap parallel workers with --slots.
Leaner calls are not always a smaller bill. On our tiny fix the checker's second read made do cost $0.012 when pinned to the chat's own model (the scene above is a separate default run, whose routed crew cost $0.066), against $0.0045 in the chat; exec, with no checker, cost $0.0032. Headless pays off when nobody is watching and a checked answer matters.
From another agent
Anything that runs a shell can hand CodeAF a job and read the result, Claude Code included. For a large, well-specified change, codeaf senior-dev --max-cost 2 --json "<brief>" works on its own branch and streams one JSON record per line. Its flags go before the brief.
Things to build
- A CI job that fixes a red build. When the test step fails, run
codeaf do "make the test suite pass" --json --timeout 20mand open a pull request only if it exits 0. - Claude Code delegating the edit. One line in
CLAUDE.md: "For implementation, runcodeaf do '<brief>' --json, read.answerand.files, and treat any exit code but 0 as not done." - A nightly triage. A cron line runs
codeaf exec "summarise what changed today and what looks risky" --jsonand posts.answerwhere your team reads. - A cost-capped big change. A script runs
codeaf senior-dev --max-cost 5 --json "<brief>"and merges its branch only if it exits 0.
Go further
senior-dev: The coding specialist, #1 on DeepSWE. Hand it one large change; it works alone on a branch of its own and checks the result itself.
Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.