Receive text as it is generated, then retrieve the saved final result. Streaming changes delivery, not model prices or token accounting. Use a backend API key with runs:write and runs:read, and a configured provider with enough credit. The API quickstart covers credentials and SDK installation.
Choose your path
| Call | Response with top-level stream: true |
|---|---|
POST /v1/inferences | HTTP 200 SSE: run identity, incremental output, saved terminal result |
| Native run, session message, customer-agent message | HTTP 202 run receipt; attach to the returned stream URL |
POST /v1/bounded-agent-runs | HTTP 202 receipt; per-model-call deltas in the durable run stream |
| Unsupported model protocol or harness | HTTP 400 streaming_not_supported, before a run or reservation is created |
Omit stream or set it to false for the existing behavior. A nested provider input.stream must agree with the top-level option. Use inferences.stream, rather than inferences.create, for incremental SDK delivery.
Anthropic Messages and OpenRouter chat support direct deltas, including raw JSON text and tool argument fragments. Codex, Claude Code, OpenCode, Hermes and Pi emit native assistant text deltas. OpenAI calls inside native harnesses retain their existing gateway streaming; there is no additional direct OpenAI endpoint.
Jev through TypeSafe/OpenRouter Decisions and the pinned DeepSeek harness reject explicit incremental streaming. Ordinary DeepSeek runs still provide committed messages and progress. Claude's dedicated structured-result feature is final-only and is not a separate streaming mode exposed by this API. Macrofold never parses partial JSON or switches models to satisfy streaming.
Direct inference in TypeScript
import { Macrofold } from 'macrofold';
const client = new Macrofold(); // MACROFOLD_API_KEY; optional baseURL for self-hosting
const request = {
model_binding: {
provider: 'anthropic' as const,
model: 'claude-haiku-4-5-20251001',
billing_mode: 'managed' as const,
},
input: {
messages: [{ role: 'user', content: 'Explain what a worktree is in two sentences.' }],
max_tokens: 256,
},
limits: { timeout_seconds: 60, max_output_tokens: 256, max_cost_micro_usd: '100000' },
};
// Persist this key if your application needs to recover after a process restart.
const key = crypto.randomUUID();
for await (const event of client.inferences.stream(request, { idempotencyKey: key })) {
if (event.type === 'run.accepted') console.log('Run:', event.run_id);
if (event.type === 'output.delta') process.stdout.write(event.data.text ?? '');
if (event.type === 'run.succeeded') console.log(event.data.result);
if (['run.failed', 'run.cancelled', 'run.timed_out'].includes(event.type)) {
console.error('Execution did not succeed:', event.data.result);
}
}
Inline stateless requests can use an unrestricted organization API key. Workspace-restricted keys must supply their authorized workspace_id. The same helper accepts typed decision definitions and explicit context; validation happens on the complete answer. For OpenRouter, select a currently enabled /v1/models route, use provider: 'openrouter', and supply its native chat body in input.
Other SDKs
All maintained helpers deliver InferenceStreamEvent objects. A terminal event contains data.result; inspect failed/cancelled/timed-out events rather than treating ordinary stream exhaustion as successful execution. HTTP/API and interrupted-transport failures raise the language's error type. Closing an iterator or stopping a callback detaches without cancelling.
| Language | Incremental helper |
|---|---|
| Python | client.inferences.stream(request_dict, request_options=options); returns a generator of Pydantic events |
| Go | client.Inferences.Stream(ctx, &request, func(event macrofold.InferenceStreamEvent) error { ... }, options...) |
| Rust | client.inferences().stream(request, |event| { ...; true }).await? |
| Java | client.inferences().stream(request, event -> { ...; return true; }) |
Each language guide includes a complete direct-streaming example.
Direct inference over REST
Set MACROFOLD_API_KEY first. Use your deployment's origin for local/self-hosted execution. Save the unique key before sending if you need recovery; retry an uncertain request with that same body and key.
export MACROFOLD_BASE_URL="https://app.macrofold.ai"
REQUEST_KEY="$(uuidgen)"
curl -N --fail-with-body "$MACROFOLD_BASE_URL/v1/inferences" \
-H "Authorization: Bearer $MACROFOLD_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-H "Idempotency-Key: $REQUEST_KEY" \
--data '{
"stream": true,
"model_binding": {
"provider": "anthropic",
"model": "claude-haiku-4-5-20251001",
"billing_mode": "managed"
},
"input": {
"messages": [{"role": "user", "content": "Explain worktrees in two sentences."}],
"max_tokens": 256
},
"limits": {"timeout_seconds": 60, "max_output_tokens": 256, "max_cost_micro_usd": "100000"}
}'
curl -N disables client buffering. HTTP 200 starts delivery; inspect the final event to determine success or failure. This example caps spending at $0.10 and requires managed credit; it does not estimate the price.
MCP does not embed a direct HTTP stream in a tool result: omit stream on direct inference tools, or use REST/SDK streaming. Native-agent MCP tools may request streaming support and return the usual receipt.
Native and bounded agents
const run = await client.runs.create({
worktree_id: process.env.MACROFOLD_WORKTREE_ID!,
agent_id: process.env.MACROFOLD_AGENT_ID!,
prompt: 'Summarize the README.',
stream: true,
});
for await (const text of client.runs.streamText(run.run_id)) process.stdout.write(text);
Use runs.events for tools, lifecycle events and model-call boundaries. Session/customer-agent messages accept the same boolean and validate the resolved session or preset harness. Native/bounded execution can queue normally. /v1/harnesses includes incremental_output only on harnesses that emit incremental assistant text; streaming alone means run-event delivery, including completed messages.
Events and final results
Direct SSE begins with run.accepted and the run ID/recovery URLs. output.delta carries text, invocation_id, message_id, content_index, and a choice_index when relevant. tool.arguments.delta carries opaque text and tool identity separately. Never execute a tool from partial arguments or concatenate independent choices into one answer. Content lifecycle/refusal events carry the same identities. Native events keep their existing harness granularity and envelope.
The terminal run.succeeded, run.failed, run.cancelled, or run.timed_out event includes the finalized result, after response persistence and financial settlement. EOF alone is not success. Missing usage stays provisional; interrupted upstream output stays incomplete. A transport.error means retrieve the result through the accepted run's URL.
One model call produces one saved response, one usage settlement and one generation trace with assembled output, subject to existing capture limits. Agent loops settle each call before the next. Reservations and permissions are checked before execution; streaming does not postpone spending controls. Direct deltas do not add per-token database or trace records. Native and bounded run events retain their existing replay storage.
Disconnects, retries and limits
Direct deltas are live-only: there are no SSE replay IDs or token replay. After disconnecting, retrieve /v1/runs/{id}/result; the run stream still provides durable lifecycle history. Reusing the same request and idempotency key while active returns HTTP 409 stream_already_started with recovery details. After completion, it returns run.accepted with delivery: result_replay followed by the saved terminal result, with no fabricated deltas or second model call. Changing the body under the same key conflicts.
SDK helpers never reconnect a direct POST after output begins. Store its run ID and idempotency key. Closing a reader does not cancel admitted work; use /v1/runs/{id}/cancel explicitly. A slow reader that exceeds the bounded delivery buffer is detached, while result processing continues. A hard process crash can lose live partial text but does not authorize repeating an ambiguous provider call.
Direct streaming requires immediate capacity: HTTP 429 streaming_capacity_unavailable rolls back admission. Retry later, or omit streaming and use Prefer: respond-async for queued acceptance. Combining stream: true with that header returns HTTP 400 invalid_request. Direct timeouts above 240 seconds return HTTP 400 streaming_timeout_unsupported; finalization has additional room within the 300-second serving limit. Native agents retain their own longer execution limits.
Direct responses retain the existing 512 KiB assembly bound. Streams send heartbeat comments while idle and request no proxy buffering. Custom hosts must supply a background lifecycle that retains the producer after disconnect; unsupported transports reject streaming before admission. Self-hosted proxies must permit a 300-second request and flush SSE promptly.