AI generation
Text generation and structured output with app.ai(), the universal method for LLM calls
Call any LLM with a single method. Pass a prompt, get text back. Pass a schema, get a validated object back. Override the model, temperature, or any other parameter per-call without changing your agent config. Works with 100+ models via LiteLLM (Python), Vercel AI SDK (TypeScript), or OpenAI-compatible APIs (Go).
from pydantic import BaseModel
class LeadScore(BaseModel):
score: float # 0-100 likelihood to convert
reasoning: str # why this score
next_action: str # recommended sales action
urgency: str # "high" | "medium" | "low"
# Structured output — pass a schema, get a validated typed object back
lead = await app.ai(
system="You are a B2B sales analyst scoring inbound leads.",
user="Company: Acme Corp, 500 employees, viewed pricing page 3x this week",
schema=LeadScore,
)
print(lead.score) # 82.5 — not a string, a float
print(lead.next_action) # "Schedule demo within 24 hours"
# Override model per-call — cheap model for fast classification
category = await app.ai(
user=ticket_text,
schema=TicketCategory,
model="openai/gpt-4o-mini", # 10x cheaper, good enough for routing
)
# Powerful model for deep analysis — same method, different model
analysis = await app.ai(
system="Provide root cause analysis with remediation steps.",
user=ticket_text,
schema=RootCauseAnalysis,
model="anthropic/claude-sonnet-4-20250514", # best reasoning for hard problems
)
# Streaming — real-time token delivery for interactive UIs
response = await app.ai(
system="Write a detailed post-mortem.",
user=incident_summary,
stream=True,
)
async for chunk in response:
print(chunk.choices[0].delta.content, end="")
# Automatic fallback chain — if primary model fails, try the next
app = Agent(
node_id="resilient-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514",
fallback_models=["openai/gpt-4o", "deepseek/deepseek-chat"],
max_cost_per_call=0.05, # hard cost cap per call
),
)What just happened
The same ai() call handled three common production modes without changing your surrounding code: typed extraction with a schema, a cheap model override for simple classification, and a stronger model override for deeper analysis. The page’s main point is that you do not need separate client stacks for each of those paths.
{
"lead_score": {
"score": 8.7,
"next_action": "schedule_demo"
},
"ticket_category": {
"category": "billing"
},
"analysis": {
"root_cause": "misconfigured webhook retry policy"
}
}
What You Get
- Text generation from 100+ LLMs via a single method call
- Structured output validated against Pydantic models (Python), Zod schemas (TypeScript), or Go structs
- Automatic prompt trimming to fit model context windows
- Rate limit retry with exponential backoff and circuit breaker
- Model fallbacks when primary model fails
- Streaming for real-time token delivery
- Multimodal input auto-detection for images, audio, and files
Patterns
System + User Prompts
The most common pattern separates instruction from input.
response = await app.ai(
system="You are a senior code reviewer. Be concise and actionable.",
user=code_diff,
)Structured Output
Schema-validated output eliminates parsing and handles malformed JSON automatically.
class RiskAssessment(BaseModel):
risk_level: str # "low", "medium", "high", "critical"
confidence: float # 0.0-1.0
findings: list[str]
recommendation: str
assessment = await app.ai(
system="Assess security risk in the following code.",
user=source_code,
schema=RiskAssessment,
model="anthropic/claude-sonnet-4-20250514",
temperature=0.0, # Deterministic for security analysis
)
if assessment.risk_level == "critical":
await escalate(assessment)Model Override Per Call
Use different models for different tasks without reconfiguring the agent.
# Fast model for classification
category = await app.ai(
user=ticket_text,
schema=TicketCategory,
model="openai/gpt-4o-mini",
)
# Powerful model for deep analysis
analysis = await app.ai(
system="Provide a thorough root cause analysis.",
user=ticket_text,
schema=RootCauseAnalysis,
model="anthropic/claude-sonnet-4-20250514",
)Streaming
Stream tokens in real-time for interactive applications.
response = await app.ai(
system="Write a detailed technical blog post.",
user="Topic: Building multi-agent systems",
stream=True,
)
# response is a LiteLLM streaming object
async for chunk in response:
print(chunk.choices[0].delta.content, end="")Model Fallbacks
Configure automatic fallback to alternative models when the primary fails.
app = Agent(
node_id="resilient-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514",
fallback_models=[
"openai/gpt-4o",
"openrouter/google/gemini-2.5-flash",
],
enable_rate_limit_retry=True,
rate_limit_max_retries=5,
),
)If the primary model hits a rate limit or returns an error, the SDK automatically tries each fallback in order.
Rate Limit Configuration
Fine-tune retry behavior for rate limits and transient errors via AIConfig.
app = Agent(
node_id="high-throughput-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514",
enable_rate_limit_retry=True,
# Retry tuning
rate_limit_max_retries=10, # max retry attempts (default: 5)
rate_limit_base_delay=1.0, # initial delay in seconds (default: 0.5)
rate_limit_max_delay=60.0, # cap on exponential backoff (default: 30.0)
rate_limit_jitter_factor=0.5, # random jitter 0.0-1.0 (default: 0.25)
# Circuit breaker — stop retrying after repeated failures
rate_limit_circuit_breaker_threshold=10, # failures before tripping (default: 5)
rate_limit_circuit_breaker_timeout=120.0, # seconds before half-open retry (default: 30)
),
)| Field | Type | Default | Description |
|---|---|---|---|
rate_limit_max_retries | int | 5 | Maximum retry attempts before failing |
rate_limit_base_delay | float | 0.5 | Initial backoff delay in seconds |
rate_limit_max_delay | float | 30.0 | Maximum backoff delay cap |
rate_limit_jitter_factor | float | 0.25 | Random jitter factor (0.0 = none, 1.0 = full) |
rate_limit_circuit_breaker_threshold | int | 5 | Consecutive failures to trip the circuit breaker |
rate_limit_circuit_breaker_timeout | int | 30 | Seconds before attempting recovery |
Cost Management
Set hard cost limits to prevent runaway spending.
app = Agent(
node_id="budget-agent",
ai_config=AIConfig(
model="anthropic/claude-sonnet-4-20250514",
max_cost_per_call=0.05, # estimated cost cap per individual call
daily_budget=10.00, # daily spend limit
),
)| Field | Type | Default | Description |
|---|---|---|---|
max_cost_per_call | float | None | Maximum estimated cost per individual AI call |
daily_budget | float | None | Maximum daily spend across all calls (resets at midnight UTC) |
Memory Scope Injection
Use memory_scope to automatically inject memory context into AI calls. The SDK fetches the specified scopes and appends their contents to the system prompt.
@app.reasoner()
async def chat(message: str) -> dict:
# Inject session memory and actor memory into the AI call
response = await app.ai(
system="You are a helpful assistant. Use the provided context.",
user=message,
memory_scope=["session", "actor"], # auto-fetches and injects memory
)
return {"response": response}
# Equivalent to manually doing:
# session_data = await app.memory.session(session_id).get("context")
# actor_data = await app.memory.actor(actor_id).get("preferences")
# then appending them to the system promptMethod Signatures
Python
async def ai(
*args, # Flexible inputs (text, URLs, bytes, dicts)
system: Optional[str] = None, # System prompt
user: Optional[str] = None, # User message (alternative to positional args)
schema: Optional[Type[BaseModel]] = None, # Pydantic model for structured output
model: Optional[str] = None, # Override model (e.g., "openai/gpt-4o")
temperature: Optional[float] = None, # Creativity (0.0-2.0)
max_tokens: Optional[int] = None, # Maximum response length
stream: Optional[bool] = None, # Enable streaming
response_format: Optional[str | Dict] = None, # "auto", "json", or "text"
memory_scope: Optional[List[str]] = None, # Memory scopes to inject
tools: Optional[...] = None, # Tool definitions (see Tool Calling page)
max_turns: Optional[int] = None, # Max LLM turns in tool-call loop
max_tool_calls: Optional[int] = None, # Max total tool calls
**kwargs, # Provider-specific parameters
) -> AnyTypeScript
AI methods are accessed via ctx inside reasoner/skill handlers:
// Inside a reasoner or skill handler:
async ctx.ai<T>(prompt: string, options?: {
system?: string; // System prompt
schema?: ZodSchema<T>; // Zod schema for structured output
model?: string; // Override model
temperature?: number; // Creativity (0.0-2.0)
maxTokens?: number; // Maximum response length
provider?: AIConfig['provider']; // Override provider
mode?: 'auto' | 'json' | 'tool'; // Structured output mode
}): Promise<T | string>
async ctx.aiStream(prompt: string, options?: AIRequestOptions): Promise<AIStream>
async ctx.aiWithTools(prompt: string, config: ToolCallConfig): Promise<ToolCallResponse>Go
func (c *Client) Complete(
ctx context.Context,
prompt string,
opts ...Option,
) (*Response, error)Options: WithSystem(s), WithModel(m), WithTemperature(t), WithMaxTokens(n), WithStream(), WithSchema(struct), WithJSONMode(), WithAPIKey(k)
SDK Reference
| Feature | Python | TypeScript | Go |
|---|---|---|---|
| Method | await app.ai(...) | await ctx.ai(...) | client.Complete(ctx, ...) |
| Schema type | Pydantic BaseModel | Zod z.object(...) | Go struct with json tags |
| Streaming | stream=True | ctx.aiStream(...) | ai.WithStream() |
| System prompt | system="..." | { system: "..." } | ai.WithSystem("...") |
| Model override | model="openai/gpt-4o" | { model: "gpt-4o" } | ai.WithModel("gpt-4o") |
| Temperature | temperature=0.7 | { temperature: 0.7 } | ai.WithTemperature(0.7) |
| Max tokens | max_tokens=1000 | { maxTokens: 1000 } | ai.WithMaxTokens(1000) |
| JSON mode | response_format="json" | { mode: 'json' } | ai.WithJSONMode() |
| Rate limit retry | Built-in (configurable) | Built-in (configurable) | Manual |
| Embeddings | Via LiteLLM | Via AIClient.embed(...) | N/A |