The open-source software factory

Solve more issues, for less.

Read the docs
apache-2.0any open modelGitHub129

CodeAFisCodeAF

3.9× as accurate as Claude Code41% cheaper than Codex2× as accurate as OpenCode6.9× cheaper than DeepSeek's harness1.5× as accurate as Pi20× as accurate as Muse Code

Run a team of agents from one window.

>● codeaf home chats sessions spend 2 want you · 11 moving · $3.42 / $50 · fri 2:46pm
────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
needs you · 2 · a digit answers it projects · folders you've opened
1 pricing: round discounts half up … @pricing · storefront ~/src/shop-api 8 chats main · 2 ahead
2 stripe webhook: retry 5 times, or 8? @hooks · billing ~/src/tinyparse 6 chats
~/src/billing 5 chats
sessions
· add rate limiting to the login route 3d spend today $3.42 of $50
· add retries to the stripe webhook 2d ▂▄▂▅▆▃▁▃▅█▆▄▆▇ 14 days $27.92 · loudest sep 21 $4.10
· add a cart item limit of 50 6d glm-5.3-flash 71% · kimi-k3 22% · qwen-max 7%
· limit checkouts to one at a time per user 1w
· add pagination to /orders 1w since you left
· rate the search results by recency 2w @cart landed with 19 passed, @lexfix landed, @vat finished
· fix the flaky test_checkout_total 2w the rate card. Two chats wait on you.
· move tinyparse to Go 1.24 3w
31 more
running · 4 · work you sent off
lexer benchmark ns/op · lexer parser 1m
invoice PDF totals · billing 3m
changelog entry for the cart fix · storefront 2m
recursive descent parser · lexer parser 40s
─ glm-5.3-flash:auto · ◇ YOLO ────────────────────────────────────────────────────────────────────── project: ~/src/shop-api ─
› type to search or start something new
→ options · opt+p project · opt+e effort · opt+a approvals · opt+k chats · / commands
>● codeaf home chats sessions spend
──────────────────────────────────────────
needs you · 2 · a digit answers it
1 pricing: round d… @pricing · storefront
2 stripe webhook: retry… @hooks · billing
sessions
· add rate limiting to the login route 3d
· add retries to the stripe webhook 2d
· add a cart item limit of 50 6d
· limit checkouts to one at a time pe… 1w
· add pagination to /orders 1w
· rate the search results by recency 2w
31 more
running · 4 · work you sent off
lexer benchmark ns/op · lexer parser 1m
invoice PDF totals · billing 3m
changelog entry for the cart fix · … 2m
recursive descent parser · lexer p… 40s
─ glm-5.3-flash:auto · ◇ YOLO ─────────── ─
› type to search or start something new
→ options · opt+p project · opt+e effort …

Start from home.

Type what you want. Home searches your past chats as you type, and enter starts a new one.

Take the tour in the docs

Alone on the Pareto frontier

Ten harnesses, one open model, the same 113 DeepSWE tasks.

tasks solved, %cost per solved task, relative to CodeAF →
0
20
40
60
1×5×10×15×20×25×

Most tasks solved on DeepSWE, at the lowest cost per solved task.

Why a factory beats a copilot

The accuracy above is not a bigger model. It is what runs underneath, and you run all of it from one window.

schemamigrationPOST /invitesexpire tokensemailsmembers pageinvite modalunit testse2emerge

From to-do list to hypergraph

Every task becomes a plan on its own: tasks nest like folders, dependencies cross them like symlinks, and every ready leaf is claimed atomically and runs at once. Hard tasks split or pivot mid-run.

fix: flaky retry test

  • workerglm-5.3-flash
  • plannerdeepseek-v4-flash
  • checkerglm-5.3-flash

outside the crew

  • reflexmistral-nemo
  • small workdeepseek-v4-flash

Every seat, Pareto-picked. Up to 25× cheaper.

Pareto Crewing matches each seat to the model that pays off for the task at hand. Frontier prices only where quality is decided, so you get frontier-level results at up to 25× lower cost.

  1. issue
  2. change

Whole harnesses, not prompts

A subharness is a full harness per job, with its own plan, checks and models. /senior-dev is the coding specialist, or build your own.

One control plane for the whole fleet

Run your whole fleet like a software factory. Every agent, every team, one window.

  • Mondays 9am: draft the weekly update
  • Always: run the tests before done
  • When the API changes: tell me

Cron, hooks and watchers, in plain English

Rules that hold, reminders on a clock, watches on the world. Say it once, scoped to this chat, this project or everywhere. No config, no DSL.

SSH-native: runs on your dev box, drive it from anywhere

The conversation lives on your dev box; desk, laptop and phone attach over your own SSH. Typing never waits on the network, and closing the lid stops nothing.

  • Held-out acceptance, like a test set

    Your request becomes a checklist the worker never sees, so it can't write tests to the list. The ending reports every point, and what no test asserts.

  • Learns from its own failures

    A failed command and the fix that worked are saved as a pair, and offered back only once proven. Below three in five, it stays silent.

  • Rewind what git can't see

    Copy-on-write forks of the full working state, .env and dev database included, in about a second. Sealed into a Merkle DAG, so any moment rewinds byte-exact.

Natively plugged into the tools your team already runs

Sign in or paste a key. Nothing to install, host or wire up.

129native connections

See all 129 connections

Open source, open models

Apache-2.0 and built in the open. If the factory unlocks something for you, a star helps more people find it.

Qwen

model: qwen

The model is the interchangeable part. CodeAF runs on the open model you point it at, and every seat of the crew can be a different one. No contract, no lock-in, no meter running against someone else's cloud.

Star on GitHub

Before you install

The questions people ask with a hand already on the install.

ask on discord
Is CodeAF a coding agent?
Yes, and then some. It splits a request into tasks, runs each on its own branch in its own copy of the repo, and merges only what passes an independent check.
Can I replace Claude Code, Codex or OpenCode with it?
Yes. Chat, edits, commands, approvals, rewind and resumed sessions are all there, plus parallel tasks, standing orders and a headless CLI. Run codeaf inside any repo.
Can I use it alongside Claude Code or Codex?
Yes. Tell them to run codeaf do "fix issue #123" --json. CodeAF plans, runs and checks the work, then hands back one JSON result and an exit code.
Which models?
Your key, your pick: OpenRouter, DeepSeek, GLM, Kimi, MiniMax, Qwen, Ollama or any OpenAI-compatible endpoint. Each kind of call can use a different model.
Does it read my AGENTS.md, CLAUDE.md and skills?
Yes. It loads the project's AGENTS.md and CLAUDE.md, and uses the skills already in your .claude, .codex, .cursor, .gemini or .agents folders where they are.
Where does it run?
One binary for macOS, Linux and Windows. codeaf chat --host devbox runs the work on a remote machine over your own SSH while the screen stays local.
What can it run without asking?
Reads, searches and git status. Everything else asks first, and destructive commands or messages sent in your name still ask, even with --yolo.
Can I undo what it did?
/rewind takes back the conversation. Attach the folder to furrow, which ships with CodeAF, and the files rewind too, even the ones git never tracked.
What stops a runaway bill?
A $500 daily backstop that pauses and asks, and any plan estimated over $100 waits for your OK. Set a tighter limit per conversation.
Does my code leave my machine?
Only to the model provider you pick; with Ollama, nothing leaves. Telemetry and the Model Pool send model-performance counts, never code, and both switch off.

Join the community

Ask a question, share what your factory shipped, and hear about new builds first.

Join the Discord

Hand the factory your next ticket.

Install CodeAF, open a project and describe the work. It asks only when it must.

Install guide