Models
100+ LLMs via LiteLLM: model-agnostic configuration, cost controls, and rate limiting
Use any LLM. Switch models per-call. Set cost caps. Auto-retry on rate limits. AgentField is model-agnostic -- the Python SDK routes through LiteLLM for 100+ models from every major provider, TypeScript uses the Vercel AI SDK, and Go uses OpenAI-compatible HTTP APIs directly.
from agentfield import Agent, AIConfig, HarnessConfig
# Agent-level config — defaults for all .ai() calls
app = Agent(
node_id="production-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514", # default model
fallback_models=[ # auto-failover chain
"openai/gpt-4o",
"deepseek/deepseek-chat",
],
max_cost_per_call=0.05, # hard cap per call
daily_budget=10.00, # daily spend limit
enable_rate_limit_retry=True, # auto-retry with backoff
rate_limit_max_retries=10,
auto_inject_memory=["user_prefs", "conversation"], # inject memory into prompts
),
# Harness uses a DIFFERENT config — coding agents have their own limits.
# No provider/model: runs AForge, the default harness, on its own model.
harness_config=HarnessConfig(
max_turns=30, # AForge bounds a run by iterations, not dollars
),
)
# Per-call model override — no agent reconfiguration needed
category = await app.ai(
user=ticket_text,
schema=TicketCategory,
model="openai/gpt-4o-mini", # cheap model for fast classification
)
analysis = await app.ai(
user=ticket_text,
schema=DeepAnalysis,
model="anthropic/claude-sonnet-4-20250514", # powerful model for reasoning
temperature=0.0, # deterministic
)
# Local models — zero data leaves your machine
local_app = Agent(
node_id="local-agent",
ai_config=AIConfig(
model="ollama/llama3",
api_base="http://localhost:11434",
),
)| Provider | Python (LiteLLM) | TypeScript (Vercel AI) | Go (HTTP) |
|---|---|---|---|
| OpenAI | openai/gpt-4o | openai provider | Default |
| Anthropic | anthropic/claude-sonnet-4-20250514 | anthropic provider | Via OpenRouter |
| Google Gemini | gemini/gemini-2.5-pro | google provider | Via OpenRouter |
| Mistral | mistral/mistral-large-latest | mistral provider | Via OpenRouter |
| DeepSeek | deepseek/deepseek-chat | deepseek provider | Via OpenRouter |
| Groq | groq/llama-3.1-70b | groq provider | Via OpenRouter |
| xAI | xai/grok-2 | xai provider | Via OpenRouter |
| Cohere | cohere/command-r-plus | cohere provider | Via OpenRouter |
| OpenRouter | openrouter/... | openrouter provider | Native |
| Ollama | ollama/llama3 | ollama provider | Native |
| Azure OpenAI | azure/gpt-4o | Via OpenAI adapter | Via URL |
| AWS Bedrock | bedrock/... | N/A | N/A |
What just happened
The example set a default model once and then overrode it only where the task changed. That is the practical pattern this page should teach: keep one baseline model for most calls, then switch to a cheaper or stronger model per execution instead of rebuilding your agent around provider-specific clients.
{
"default_model": "gpt-4o",
"classification_override": "gpt-4o-mini",
"analysis_override": "anthropic/claude-sonnet-4-20250514"
}
Configuration Examples
Minimal Setup
from agentfield import Agent, AIConfig
# Uses OPENAI_API_KEY from environment
app = Agent(
node_id="my-agent",
ai_config=AIConfig(model="openai/gpt-4o"),
)Cost-Conscious Production
app = Agent(
node_id="production-agent",
ai_config=AIConfig(
model="openai/gpt-4o-mini",
max_cost_per_call=0.05,
daily_budget=10.00,
fallback_models=["openrouter/google/gemini-2.5-flash"],
enable_rate_limit_retry=True,
rate_limit_max_retries=10,
max_tokens=2000,
),
)Multi-Provider Resilience
app = Agent(
node_id="resilient-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514",
fallback_models=[
"openai/gpt-4o",
"openrouter/google/gemini-2.5-pro",
"deepseek/deepseek-chat",
],
timeout=30,
retry_attempts=3,
),
)Local Models with Ollama
app = Agent(
node_id="local-agent",
ai_config=AIConfig(
model="ollama/llama3",
api_base="http://localhost:11434",
),
)TypeScript with Anthropic
const app = new Agent({
nodeId: 'my-agent',
aiConfig: {
provider: 'anthropic',
model: 'claude-sonnet-4-20250514',
apiKey: process.env.ANTHROPIC_API_KEY,
temperature: 0.3,
maxTokens: 4096,
},
});Go with OpenRouter
config := &ai.Config{
APIKey: os.Getenv("OPENROUTER_API_KEY"),
BaseURL: "https://openrouter.ai/api/v1",
Model: "anthropic/claude-sonnet-4-20250514",
}
client, _ := ai.NewClient(config)Model Selection Guide
| Use Case | Recommended | Why |
|---|---|---|
| General tasks | openai/gpt-4o | Good balance of capability and cost |
| Fast classification | openai/gpt-4o-mini | 10x cheaper, fast |
| Deep reasoning | anthropic/claude-sonnet-4-20250514 | Best for code and analysis |
| Long context | openrouter/google/gemini-2.5-pro | 2M token context window |
| Budget-friendly | deepseek/deepseek-chat | Strong capability at low cost |
| Local/private | ollama/llama3 | No data leaves your machine |
| Maximum quality | anthropic/claude-opus-4-20250514 | Best reasoning, highest cost |
AIConfig (Python)
The AIConfig class controls all LLM behavior. Set defaults at agent construction time, override per-call.
Core Fields
| Field | Type | Default | Description |
|---|---|---|---|
model | str | "gpt-4o" | Default model (use provider/model format for LiteLLM) |
temperature | float? | None | Creativity (0.0-2.0). None = model default |
max_tokens | int? | None | Maximum response tokens. None = model default |
top_p | float? | None | Nucleus sampling (0.0-1.0). None = model default |
stream | bool? | None | Enable streaming. None = model default |
response_format | str | "auto" | "auto", "json", or "text" |
API Configuration
| Field | Type | Default | Description |
|---|---|---|---|
api_key | str? | None | API key (overrides env vars) |
api_base | str? | None | Custom API base URL |
api_version | str? | None | API version (for Azure) |
organization | str? | None | Organization ID (for OpenAI) |
litellm_params | dict | {} | Additional LiteLLM parameters |
Multimodal Models
| Field | Type | Default | Description |
|---|---|---|---|
vision_model | str | "dall-e-3" | Model for image generation |
audio_model | str | "tts-1" | Model for speech generation |
video_model | str | "fal-ai/minimax-video/image-to-video" | Model for video generation |
image_quality | str | "high" | "low" or "high" |
audio_format | str | "wav" | Default audio format |
fal_api_key | str? | None | Fal.ai API key (or FAL_KEY env var) |
Cost Controls
| Field | Type | Default | Description |
|---|---|---|---|
max_cost_per_call | float? | None | Maximum cost per AI call in USD |
daily_budget | float? | None | Daily budget for AI calls in USD |
Rate Limiting
| Field | Type | Default | Description |
|---|---|---|---|
enable_rate_limit_retry | bool | True | Auto-retry on rate limit errors |
rate_limit_max_retries | int | 5 | Maximum retry attempts |
rate_limit_base_delay | float | 0.5 | Base delay for exponential backoff (seconds) |
rate_limit_max_delay | float | 30.0 | Maximum backoff delay (seconds) |
rate_limit_jitter_factor | float | 0.25 | Randomization factor (+/- 25%) |
rate_limit_circuit_breaker_threshold | int | 5 | Consecutive failures before circuit opens |
rate_limit_circuit_breaker_timeout | int | 30 | Circuit breaker timeout (seconds) |
Resilience
| Field | Type | Default | Description |
|---|---|---|---|
fallback_models | list[str] | [] | Models to try if primary fails |
timeout | int? | None | Call timeout in seconds |
retry_attempts | int? | None | Retry attempts for failed calls |
retry_delay | float | 1.0 | Delay between retries (seconds) |
Context Management
| Field | Type | Default | Description |
|---|---|---|---|
max_input_tokens | int? | None | Max input tokens (overrides auto-detection) |
preserve_context | bool | True | Preserve conversation context |
context_window | int | 10 | Previous messages to include |
auto_inject_memory | list[str] | [] | Memory scopes to auto-inject |
AIConfig (TypeScript)
The TypeScript SDK uses the Vercel AI SDK and requires an explicit provider.
| Field | Type | Default | Description |
|---|---|---|---|
provider | string | "openai" | "openai", "anthropic", "google", "mistral", "groq", "xai", "deepseek", "cohere", "openrouter", "ollama" |
model | string? | "gpt-4o" | Model name |
embeddingModel | string? | None | Embedding model |
apiKey | string? | None | API key |
baseUrl | string? | None | Custom base URL |
temperature | number? | None | Creativity (0.0-2.0) |
maxTokens | number? | None | Maximum response tokens |
enableRateLimitRetry | bool | true | Auto-retry on rate limits |
rateLimitMaxRetries | number | 20 | Maximum retry attempts |
rateLimitBaseDelay | number | 1.0 | Base delay (seconds) |
rateLimitMaxDelay | number | 300.0 | Maximum delay (seconds) |
rateLimitJitterFactor | number | 0.25 | Randomization factor (+/- 25%) |
rateLimitCircuitBreakerThreshold | number | 10 | Consecutive failures before circuit opens |
rateLimitCircuitBreakerTimeout | number | 300 | Circuit breaker timeout (seconds) |
Config (Go)
The Go SDK uses a simple config struct for OpenAI-compatible APIs.
| Field | Type | Default | Description |
|---|---|---|---|
APIKey | string | OPENAI_API_KEY env | API key |
BaseURL | string | https://api.openai.com/v1 | API endpoint |
Model | string | "gpt-4o" | Default model |
Temperature | float64 | 0.7 | Creativity |
MaxTokens | int | 4096 | Maximum tokens |
Timeout | time.Duration | 30s | HTTP timeout |
SiteURL | string | "" | OpenRouter site URL |
SiteName | string | "" | OpenRouter site name |
The Go SDK auto-detects OpenRouter from the OPENROUTER_API_KEY environment variable.
Environment Variables
Set API keys via environment variables. The SDK picks them up automatically.
| Variable | Provider |
|---|---|
OPENAI_API_KEY | OpenAI |
ANTHROPIC_API_KEY | Anthropic |
GOOGLE_API_KEY | Google Gemini |
MISTRAL_API_KEY | Mistral |
DEEPSEEK_API_KEY | DeepSeek |
GROQ_API_KEY | Groq |
XAI_API_KEY | xAI |
COHERE_API_KEY | Cohere |
OPENROUTER_API_KEY | OpenRouter |
AZURE_OPENAI_API_KEY | Azure OpenAI |
FAL_KEY | Fal.ai (media generation) |