curl example, choose your Cloud region. For authenticated calls, it must be the same region as your key. The examples below use MIROBODY_API_BASE:
Endpoint
- Client function tools — inject your own tools; the model hands off via
function_calloutput items (Function calling). - Stored conversations —
storedefaults totrue; chain turns withprevious_response_idor bind a durable conversation withsession_id(State & memory). - Standard
response.*streaming events (Streaming).
/v1 endpoints share one base URL — pick the cluster your account uses:
/v1 — it serves its own /api/* routes and an /mcp endpoint, described in the Self-Host Mirobody.
Japan and EU clusters are in preparation — see Regions. Model providers and pricing can differ by region, so read GET /v1/models from the cluster you call.
Create a response
Request body
The
user field is the tenant-isolation key. The backend maps (your account, user) to an internal Subject, and Subjects are fully isolated from one another — pass each end-user’s stable id and no user ever sees another’s data. Omit it and the call falls back to your account’s default Subject. Subjects are invisible to the Mirobody web app and to other developers.The value is normalized (trimmed + lowercased) before mapping, so Alice and alice resolve to the same Subject. Pass a stable, canonical id.Response object
output array contains only standard OpenAI item types — reasoning, message, and (on a client-tool handoff) function_call. Official SDKs parse it as-is.
Server-side built-in tool runs are deliberately not output items. The trace rides in the top-level tool_steps extension field ({id, name, arguments, result} — the same shape as the Answers API), which SDKs safely ignore. In streaming it surfaces as the side-channel event response.mirobody_tool_call. health_records and citations carry the evidence the answer used, exactly as on the Answers API.
Usage accounting
usage.input_tokens reports the input you actually sent; the platform’s system prompt / tool-schema overhead is broken out as input_tokens_details.system_tokens. output_tokens_details.reasoning_tokens reports provider reasoning tokens. Counts are summed across every model call of the agent turn.
usage.billed_tokens ({input, output, total}) is what you’re metered on — present on both the Agent API and the Answers API. Applying the catalog’s standard input/output rates gives an estimate; prompt-cache discounts can make actual metered cost lower.
Retrieve a stored response
store=true responses are retrievable; expired (30-day TTL) or deleted responses return 404.
Delete a stored response
State model
store and retention are orthogonal — one governs the conversation, the other governs the data plane:
retention governs explicit data-plane writes; it does not govern a stored conversation. A stored conversation can also produce health records and memories from its content — delete those with DELETE /v1/data, or erase the Subject. store=false conversations produce none.
Full details — TTLs, chaining semantics, stateless replay, and the cross-session memory that store=true feeds — in State & memory.
Strict validation
By default the surface is lenient (compatibility-first) — an unknown top-level parameter is ignored. Setstrict: true in the body or send the header X-Mirobody-Strict: 1 to make it a hard 400 unsupported_parameter instead, with param pointing at the first offending key. Both surfaces (/v1/responses and /v1/chat/completions) honor it. Turn it on while integrating to catch typos and misplaced fields early.
reasoning (e.g. {"effort": "low" | "medium" | "high"}) is honored: the effort maps to the underlying model’s extended-thinking / reasoning budget. It streams as the response.reasoning_summary_text.* events and lands in the non-stream reasoning output item; usage.output_tokens_details.reasoning_tokens counts it. Omit reasoning for the model’s default behavior. Models without a thinking mode ignore it (no error).Errors
Standard error envelope. Surface-specific cases:See also
- Function Calling — declaring client tools and continuing after a handoff.
- State & Memory —
store, session state and cross-session memory. - Streaming — the
response.*event sequence. - Backbone Mode — turning this endpoint into a bare inference backend.