Harness
Run a coding agent from inside agent loops with turn caps, schema-bound output, and retries. AForge is the default; Claude Code, Codex, Gemini CLI, and OpenCode are one field away.
The building block for driving a coding agent from inside your loop. app.harness(prompt, schema=...) runs AForge, AgentField's native harness, by default — or Claude Code, Codex, Gemini CLI, or OpenCode when you choose them. Either way it waits for structured output, validates it against your schema, and hands back a rich result with cost and retry metadata.
AForge is provisioned alongside the af binary — by the curl installer, AgentField Desktop, and the official agent Docker images — so the default path has nothing to install. Export OPENROUTER_API_KEY and the call below works.
Without harness, you would manage subprocess plumbing, JSON repair, retry-with-backoff, schema validation, output-file cleanup, and per-provider quirks yourself. With harness, those become one call you can wrap in a loop, a tournament, or an adversarial pattern.
The shape
from pydantic import BaseModel
from agentfield import Agent, HarnessConfig
class ReviewResult(BaseModel):
findings: list[str]
severity: str # "low" | "medium" | "high"
# Defaults pinned on the agent — every .harness(...) call inherits them.
# No provider, no model: this runs on AForge with its own default model.
app = Agent(
node_id="reviewer",
harness_config=HarnessConfig(
max_turns=12,
),
)
@app.reasoner()
async def review_diff(diff: str) -> dict:
result = await app.harness(
f"Review this diff. Be precise, no fluff.\n\n{diff}",
schema=ReviewResult,
)
if result.is_error:
return {"ok": False, "error": result.error_message}
return {
"ok": True,
"review": result.parsed.model_dump(),
"cost_usd": result.cost_usd,
"num_turns": result.num_turns,
"session_id": result.session_id,
}Choose your worker
The default is a starting point, not a lock-in. Set provider and the same call runs on a
different coding agent — that is how you orchestrate Claude Code, Codex, Gemini CLI, or
OpenCode from inside an AgentField loop, or run a fleet of them side by side.
result = await app.harness(prompt, schema=ReviewResult, provider="claude-code")
# Claude Code is also the provider that enforces a USD cost cap and a tool allowlist.
capped = await app.harness(
prompt,
schema=ReviewResult,
provider="claude-code",
permission_mode="plan",
max_budget_usd=0.50,
tools=["Read", "Grep", "Glob"],
)The USD cost cap and the tool allowlist are claude-code only — AForge ignores both, so bound a default-provider run with max_turns instead.
Each override needs its own worker installed and its own credential. The default needs no install — only OPENROUTER_API_KEY.
provider | Install | Credential | Best at |
|---|---|---|---|
aforge (default) | ships with af — nothing to do | OPENROUTER_API_KEY | the zero-setup path; docs |
claude-code | Python: pip install 'agentfield[harness-claude]' · TypeScript: npm install @anthropic-ai/claude-agent-sdk · Go: npm install -g @anthropic-ai/claude-code | ANTHROPIC_API_KEY | the careful reasoner |
codex | npm install -g @openai/codex | OPENAI_API_KEY | the fast implementer |
gemini | npm install -g @google/gemini-cli | GEMINI_API_KEY | the long-context worker |
opencode | curl -fsSL https://opencode.ai/install | bash | opencode auth login | the open-model path |
Provider selection resolves in this order, first match wins:
- An explicit
provideron the call, then on the agent'sHarnessConfig. - The
AGENTFIELD_HARNESS_PROVIDERenvironment variable —aforge,claude-code,codex,gemini, oropencode(the Python SDK also acceptsgrok). aforge.
HarnessConfig — defaults on the Agent
Pin defaults on the agent constructor; override per call only for the dimensions that genuinely vary by task.
| Field | Type | Default | What it does |
|---|---|---|---|
provider | string | "aforge" | "aforge" / "claude-code" / "codex" / "gemini" / "opencode". Leave unset for AForge; AGENTFIELD_HARNESS_PROVIDER shifts the default. |
model | string | empty | Empty means the provider's own default (AForge uses its default, overridable with AFORGE_MODEL). Set it to pin a provider-specific identifier — e.g. "sonnet", "gpt-5-codex", "gemini-2.5-pro", "qwen/qwen3-coder". |
max_turns | int | 30 | Hard cap on agent iterations. |
max_budget_usd | float | null | USD cost cap — enforced by claude-code only. AForge and the other CLI providers ignore it (AForge's own --budget is a token budget, not a USD cap); bound those runs with max_turns instead. |
max_retries | int | 3 | Retry attempts for transient errors (rate-limit, 5xx, connection reset). |
initial_delay / max_delay / backoff_factor | float | 1.0 / 30.0 / 2.0 | Exponential-backoff knobs for retries. |
tools | string[] | ["Read","Write","Edit","Bash","Glob","Grep"] | Allowed tool names — claude-code only. AForge and the other CLI providers ignore the list and run with their own toolset. |
permission_mode | string | null | "plan" (plan-first, then execute) or "auto" (bypass per-step prompts). Honoured by claude-code, codex, and gemini; ignored by the default aforge and by opencode. |
system_prompt | string | null | Custom system prompt prepended to the loop. |
env | dict | {} | Extra environment variables forwarded to the subprocess. |
cwd | string | working dir | Working directory the coding agent treats as the repo root. |
project_dir | string | null | opencode only — maps to --dir. When set, cwd is used only for output-file placement. |
aforge_bin / codex_bin / gemini_bin / opencode_bin | string | binary name | Override CLI paths when the binary is not on $PATH — every provider binary is resolved by a plain $PATH lookup (the Go SDK uses a single BinPath). AForge also reads the AFORGE_BIN env var; af aforge ensure (re)installs it into $AGENTFIELD_HOME/bin (default ~/.agentfield/bin). |
HarnessResult — what comes back
| Field | Type | What it is |
|---|---|---|
result / text | string | Raw agent response (last message text). |
parsed | model / null | Validated schema instance when schema= was passed. null if validation fell through all repair layers. |
is_error / error_message | bool / string | True when the run failed terminally. Inspect error_message for the diagnosis. |
cost_usd | float / null | Total cost reported by the provider. null when the provider does not surface cost. |
num_turns | int | Iterations the agent took. Useful for cost monitoring and tuning max_turns. |
failure_type | string | none, crash, timeout, api_error, no_output, or schema (failureType in TypeScript). See Error handling. |
duration_ms | int | Wall-clock execution time. |
session_id | string | Provider session identifier — pass back in resume_session_id to continue a multi-turn run. |
messages | list | Full message stream (provider-specific shape). Inspect for debugging. |
Schema-bound output
Pass a Pydantic model (Python), Zod schema (TypeScript), or Go struct as schema=. The runner injects an OUTPUT REQUIREMENTS suffix telling the agent to write JSON to .agentfield_output.json, then reads, repairs, validates, and returns result.parsed as a validated instance. Three recovery layers run before declaring failure:
- Parse the output file directly.
- Cosmetic repair — strip markdown fences, trailing commas, repair truncated braces.
- One-shot AI repair — re-emit the same content as valid JSON conforming to the schema (no tools, no exploration; cheap reformatting only).
After that, the run is retried up to max_retries times. The output file is cleaned up automatically.
The agent writes its answer to a file with its own Write tool instead of the harness parsing JSON out of response text. Every provider has a Write tool, so the same strategy works on all of them:
Provider switching is a one-field flip
Every provider is interchangeable through provider=. Same loop code, different worker:
# Plan with Claude (careful reasoner), execute with Codex (fast implementer).
plan = await app.harness(prompt, provider="claude-code", permission_mode="plan", schema=ChangePlan)
edits = await app.harness(apply, provider="codex", permission_mode="auto", schema=EditReport)See the provider docs for the role each one plays best:
- AForge: the default. Ships with
af, nothing to install. - Claude Code: the careful reasoner.
- Codex: the fast implementer.
- Gemini CLI: the long-context worker.
- OpenCode: the open-model path.
Error handling
A failed run does not raise. Check is_error, then branch on failure_type to tell the failure modes apart:
result = await app.harness("Deploy the staging environment.", max_turns=20)
if result.is_error:
match result.failure_type:
case "timeout":
log.warning(f"Harness timed out after {result.duration_ms}ms")
case "crash":
log.error(f"Agent crashed: {result.error_message}")
case "schema":
log.warning("Output did not match schema after retries")
case "api_error":
log.error("Transient API error, will retry on next invocation")
else:
print(result.text)
In Go, compare result.FailureType with harness.FailureTimeout, harness.FailureCrash and harness.FailureSchema.
Common patterns
Retry with provider fallback
When the primary provider trips a budget or rate-limit, fall back to another worker without losing the loop's intent. Starting from the default keeps the happy path install-free.
async def hardened_harness(prompt: str, schema, **opts):
for provider in ("aforge", "claude-code"):
result = await app.harness(prompt, schema=schema, provider=provider, **opts)
if not result.is_error and result.parsed is not None:
return result
raise RuntimeError(f"all providers failed: {result.error_message}")Tournament
Run multiple providers in parallel against the same prompt and pick the best result with an LLM-as-judge call.
runs = await asyncio.gather(*[
app.harness(prompt, provider=p, schema=ImplResult, max_budget_usd=0.05)
for p in ("claude-code", "codex", "opencode")
])
verdict = await app.harness(
"Pick the best implementation. Score on correctness and minimality.\n\n"
+ format_runs(runs),
provider="claude-code",
schema=Verdict,
)
return runs[verdict.parsed.winner_index]Adversarial verifier
One provider proposes a change, a second one verifies it. The verifier's output triggers re-runs until severity drops or the budget caps out.
for attempt in range(3):
impl = await app.harness(prompt, provider="codex", schema=ImplResult)
if impl.is_error:
continue
review = await app.harness(
f"Find real bugs in this change:\n\n{impl.parsed.diff}",
provider="claude-code",
permission_mode="plan",
schema=ReviewResult,
)
if review.parsed.severity in ("low", "medium"):
return impl # accepted
prompt = f"Re-attempt. Previous reviewer findings:\n{review.parsed.findings}\n\n{prompt}"Authentication
Each provider reads its own credentials from the environment forwarded to the harness subprocess. None of these need to be passed through harness_config.env if they're already in the process environment.
| Provider | Required env vars |
|---|---|
aforge (default) | OPENROUTER_API_KEY. AFORGE_MODEL optionally overrides the model. |
claude-code | ANTHROPIC_API_KEY (or Vertex / Bedrock routing via claude_agent_sdk config) |
codex | OPENAI_API_KEY, or run codex login once for OAuth |
gemini | GEMINI_API_KEY or GOOGLE_API_KEY, or run gemini once to sign in interactively |
opencode | Whatever its opencode auth login flow configures — OPENROUTER_API_KEY, OLLAMA_HOST, etc. |
Check what a provider can actually see before you run a loop against it:
af harness doctor --provider aforge # or claude-code / codex / gemini / opencode
See also
- Cross-agent calls — call other AgentField nodes from inside a harness loop, or vice versa.
- AForge: the default harness
- Claude Code
- Codex
- Gemini CLI
- OpenCode