Skip to main content
October 8, 2026 • External API Update The first message of a new session used to wait for two things before the first token: the one-time build of your organization’s agents, and the orchestrator’s routing call that decides which agent should answer. This release lets integrators remove both from the critical path.

target_agent_id on Chat Completions

When your app already knows which agent should answer, pass its id and the orchestrator’s routing LLM call is skipped on that turn. The target agent receives everything the orchestrator would have received — the user’s memory and your extra_context. Later turns continue with that agent automatically. Requires the agents:invoke scope (default keys have it). → Chat Completions: Low-latency first turn

POST /v1/warmup

Build the orchestrator for a user — and optionally one agent — right after provisioning, with no LLM call and no charge. The first message then starts immediately. → Warmup

extra_context reaches the answering agent

Per-request context is now delivered to the agent that actually answers (the routed sub-agent, the target_agent_id target, or the agent called through /agents/{agent_id}/completions), not only to the orchestrator. It stays a non-persistent trailing block: never stored in the session, never merged into the system instruction.

Streaming keepalives during builds

Streaming responses now send SSE comment lines (: keepalive) every 15 seconds while the server is still building agents or waiting on tools, so load balancers and HTTP clients with idle timeouts never see a silent connection. Standard SSE clients ignore comment lines.