Skip to the content.

Inference Runtime Guide

core/inference is the unified, instance-owned runtime for model inference. Active workloads are Generate, Embed, and Transcription; they are enumerated once, by model.Operations() in core/inference/model.

The model declaration vocabulary — identity, descriptor, capabilities, limits, and lifecycle — lives in core/inference/model. core/inference re-exports it through deprecated aliases (inference.ModelRef and friends) so existing callers keep compiling; new code should import core/inference/model directly.

Exact addressing

Every call takes a concrete ModelRef:

model := model.ModelRef{
    ID: model.ModelID{
        Provider: "deepseek",
        Name:     "deepseek-flash",
    },
}

The runtime never picks or replaces a model. Optional routing is provided by core/inference/route. Generate routing consults the providers' declared model capabilities on both the selection and the fallback path: targets whose declared output kinds cannot serve the request intent are skipped, while models with undeclared capabilities are treated as undeclared (not unsupported) — preflight remains the final arbiter for those.

Deployment config

Providers and the assembly are separate resources:

resources:
  provider:
    kind: inference.Provider
    impl: openai
    settings:
      id: deepseek
      spec:
        api: responses
        endpoint:
          base_url: https://api.deepseek.com
        wire:
          reasoning_channel: text   # DeepSeek streams plain reasoning text
        models:
          - name: deepseek-flash
            kind: generate
            capabilities:
              inputs: [text, image, data, tool_call, tool_result]
              outputs: [text]
      profiles:
        - secrets:
            api_key: ${env:DEEPSEEK_API_KEY}
  infer:
    kind: inference.Assembly
    impl: unified
    deps:
      provider: provider

Provider implementations are registered by the application from provider driver modules:

reg.MustRegister(openai.Factory())
reg.MustRegister(inference.Factory{})

The provider spec is layered — endpoint (where the API is and how the key rides), wire (which dialect the endpoint speaks), models (the line-up this deployment serves) — and core/inference treats every provider through the same declaration vocabulary. The full key reference, including wire.store (false by default, or "omit" to keep the retention field off the request entirely for endpoints whose schema does not know it) and wire.video_input (a compatible-endpoint extension that, together with a model declaring video input, lets a chat-surface deployment carry video parts), lives in the flowcraft-config provider reference.

Provider lifecycle

Opening a model resolves its credentials and constructs its provider clients. The default call path (Assembly.Generate, Embed, Transcribe, ...) resolves the target per call, so nothing is cached behind the caller's back and a deployment never needs credentials at build time: Validate and InspectModel work without them, and a missing credential surfaces on first use.

Callers that want opened drivers to outlive a single call say so explicitly. Assembly.Bind(ctx, ref) opens every operation the model declares and returns a Binding: immutable, safe for concurrent use, owning its drivers until dropped. Binding again picks up rotated credentials. A binding compiles per request:

binding, err := assembly.Bind(ctx, model)
prepared, err := binding.PrepareGenerate(ctx, request) // compiles once
response, err := prepared.Execute(ctx)                 // provider I/O only

A Prepared attempt is what keeps preflight and execution from compiling the same request twice. Assembly.Prepare* returns the same kind of handle for callers that do not hold a binding, and executing one performs provider I/O only while still carrying the span, metrics, and usage envelope of a direct call. Routing uses exactly this: a routed attempt opens and compiles once, then executes that compilation.

The graph inference node is the in-tree example of the binding lifetime: a node's model is static graph config and the node outlives the turns that run through it, so it opens each configured model once (keyed by the full model reference, shared by every node in the graph) and compiles per turn.

The script bridge does the same for script calls: inference.generate, explain, stream, embed, transcribe, and transcribeSession resolve their model through one inference.BindingCache owned by the bridge, so a script that keeps addressing the same model reuses its drivers. Both hosts rely on the same cache type, and both keep working when the model is not one they have seen before: a miss simply opens it.

That reuse has a staleness window: a credential or client setting rotated after a model was opened stays in effect until the deployment is rebuilt (or the node addresses a different profile). Opening is logged with llm.provider, llm.model, and llm.profile, so "the deployment is still using the previous key" is diagnosable rather than silent. The node keeps at most 16 bindings; past that it evicts the least recently used one, so a board-derived model reference cannot accumulate drivers while the models the graph keeps addressing stay open.

Model declarations

Every published model comes from the provider spec's models list: no driver ships a line-up, so a model exists because the deployment declares it and the declaration is the whole fact. Nothing is inherited from a same-named model, which is why a generate model has to state its text output and every published capability has to be written out; the catalog key that used to select a namespace is retired and rejected with the migration path. Driver control facts that no capability kind expresses are declared per model too (Bytedance's max_resolution and video parameter matrix, MiniMax's wire_model and video surface).

Capability declarations are promises validated per provider surface: a model published with reasoning kind toggle must compile reasoning_enabled=false on that surface, and one published as always rejects it. Discovery bits such as hosted_web_search and custom_embed_dimensions ride on the model descriptor, so hosts can surface per-model options without driver-specific knowledge.

Declared limits are enforced before any provider work. A Generate request whose max_output_tokens exceeds the target model's declared max_output_tokens is rejected at declaration time — the driver is not opened, and the rejection is transport-safe, so a routed request falls back to a target that declares room for it instead of failing. The limit is a promise, not a clamp: the caller's budget is never silently lowered, and an undeclared limit rejects nothing.

max_input_tokens is not enforced: the module has no tokenizer, and a missing or approximate count would be worse than none. Hosts that own the conversation history can read the declared input window from InspectModel and enforce it where they already trim context.

Reasoning provenance

Reasoning models return a trace — thinking text plus a signature or an encrypted payload — that providers expect back in later turns, and they verify it against the model and account that produced it. OpenAI validates a reasoning item's id and encrypted content; Anthropic validates a thinking block's signature. A payload minted elsewhere is rejected with a provider 400 that no retry and no fallback repairs, because the trace stays in the conversation.

Drivers therefore stamp every trace they produce with a verification scope (message.ReasoningPart.Source) and replay a stored trace only when the target's scope matches:

Trace Compile outcome
stamped with this target's scope replayed, subject to the surface's own rules (id + encrypted payload, signature, ...)
stamped with another scope dropped, with both scopes named in the compile report
unstamped (stored before drivers stamped, or built by hand) dropped as unattributable

The derived scope is the address that produced the trace — provider, model, and the credential profile when the reference names one:

openai/gpt-5.6-luna/default

Switching models inside one deployment therefore drops the previous model's traces. That is what Anthropic requires (a thinking signature is bound to the model), and it is what an endpoint that does not document cross-model verification cannot be trusted with. A deployment that has verified a set of models or credentials accept each other's traces declares one shared scope and every model of that deployment uses it:

spec:
  wire:
    reasoning_scope: openai-prod-shared   # OpenAI and Anthropic drivers

Bytedance uses the same key at the top level of spec (reasoning_scope), because its spec is flat. Ark consumes no reasoning input today, so the stamp only keeps its traces attributable when a conversation moves elsewhere.

Upgrading: a transcript stored before this rule carries unstamped traces, so those traces are dropped instead of replayed. The assistant text is untouched — only thinking continuity is lost — and the conversation no longer risks a provider rejection it cannot recover from.

Routing

Optional target selection is an inference.Router resource. It consumes one inference.Assembly as its target dep and reads the route policy from its own settings:

resources:
  router:
    kind: inference.Router
    impl: unified
    deps:
      target: infer
    settings:
      generate:
        - tier: fast
          targets:
            - model: {id: {provider: deepseek, name: deepseek-flash}}
              score: {quality: 0.8, speed: 0.9}
      retry:
        generate:
          max_attempts: 2
          max_total_attempts: 5
          backoff:
            kind: exponential   # fixed | exponential (default)
            initial: 100ms
            max: 2s
            multiplier: 2
            jitter: full        # none | equal | full (default)
          retryable: [rate_limit, timeout, unavailable]
          fallback_on_retry_exhausted: true
      circuit_breaker:
        failure_threshold: 5      # consecutive transient failures; default 5
        recovery_window: 30s      # open-circuit window; default 30s
        half_open_max_probes: 1   # concurrent probes while half-open; default 1

The policy has three operation areas — generate, embed, and transcription — each a list of tier pools. A pool is an allowlist of exact model targets plus optional normalized score signals (quality / economy / speed / reliability, all in [0, 1]). Scores guide selection only; they never claim a request is executable.

Request metadata

GenerateRequest.RequestMetadata is an opaque map[string]string carried with every call. Core never interprets its keys; callers and hosts decide the vocabulary (conversation identifiers, turn identifiers, installation metadata, ...). Graph inference nodes expose the same field as request_metadata node config, and script/direct callers can set it on the canonical request.

Drivers forward the bag only when their deployment enables it, because each provider API has a different legal shape. Provider specs accept:

settings:
  spec:
    request_metadata:
      envelope: metadata          # any non-empty top-level field name; empty disables

OpenAI-compatible drivers map metadata onto the native OpenAI metadata object; client_metadata is emitted as a passthrough object for gateways that speak the Codex convention. An empty configuration never sends anything, and core keys are forwarded verbatim.

request_metadata forwarding is implemented by the OpenAI driver, which covers OpenAI, Azure, DeepSeek, Kimi and any compatible endpoint configured through its endpoint block. Anthropic, MiniMax, and Bytedance are the current exceptions: their official Messages/Ark surfaces do not model arbitrary request metadata and the drivers deliberately keep their native transport paths, so canonical metadata is not forwarded until those SDKs/providers add a native channel.

Because RequestMetadata is part of the compile ledger, drivers that cannot forward it report a dropped decision instead of silently ignoring it. Drivers that forward it report native when their deployment enables an envelope. The envelope is an arbitrary non-empty string naming the top-level body field; providers that type metadata natively lower that name through their SDK types, while other names ride as passthrough JSON fields.

Unmodeled provider fields (json_set)

Compatible endpoints extend the OpenAI schema with knobs no SDK models — Kimi's thinking, Qwen's enable_thinking / thinking_budget, a gateway's own routing field. The OpenAI driver carries those as an explicit per-request extension no matter which deployment id serves the call:

{
  "provider": "kimi",
  "extension": "generate_options",
  "value": {
    "json_set": {
      "enable_thinking": false,
      "thinking.keep": "all"
    }
  }
}

Each key is a path in sjson notation (a dot descends into an object, so thinking.keep sets one leaf and leaves its siblings alone) and each value is the raw JSON to place there. The compile report names every key it applied, so the ledger says "this value rode the request" — and nothing more:

Key Behavior
model, messages, input, tools, tool_choice, response_format, text, max_tokens / max_completion_tokens / max_output_tokens, temperature, top_p, n, store, metadata, reasoning / reasoning_effort, service_tier, verbosity, parallel_tool_calls, max_tool_calls, safety_identifier, prompt_cache_key, modalities, audio rejected: the compiler lowers these from the canonical request, and a patch would make the report claim a decision the body contradicts
anything else carried verbatim, at most 32 keys and 64 KiB per request

This is deliberate passthrough, not a capability claim: FlowCraft cannot validate what an endpoint does with a field it did not model, so routing and preflight never learn anything from json_set. A request that needs the endpoint to do something with a trace or a content kind — reasoning round-trips, video input — needs the driver-side support, not a body patch.

json_set applies to the JSON generate surfaces (Responses and Chat Completions, unary and stream). Multipart transports do not run it.

Deployment defaults: wire.extra_body

An endpoint whose dialect is fixed — "this GLM instance always thinks", "this gateway always wants its own routing field" — declares the same shape in the provider spec, so every request the deployment serves carries it:

settings:
  spec:
    api: chat
    wire:
      extra_body:
        thinking: {type: enabled}

Keys and values are the json_set vocabulary (sjson paths, raw JSON), with the same bounds (32 keys / 64 KiB) and the same rejected keys. The two sources compose in a fixed order: wire.extra_body is written first, then the request's json_set, so a request key wins — an identical path replaces the deployment value, a nested path updates the object the deployment wrote.

The difference between the two is scope and reporting, not shape: extra_body is deployment configuration and, like store or endpoint.headers, carries no per-request decision; json_set is a request decision and appears in the compile report key by key. Both reach the generate surfaces only, so a deployment whose declared line-up has no generate model is rejected at build time rather than carrying a field that can never apply.

Component notes

Decisions are field-level: every active canonical field carries exactly one terminal disposition. Some fields aggregate several content parts under one path — the * in generate.context.*.content.parts.tool_result spans every message — so a single disposition cannot say "the text arrived but the image did not". Such a decision may carry components notes: one entry per part, in encounter order, each naming the content kind, its position, and its own disposition. The notes are populated only when the field lost at least one part, and they must fold onto the field's disposition, so a field can never read native while one of its components was dropped.

The OpenAI driver uses them for multimodal tool results: text that rides along is native, an image the wire cannot carry (a Chat Completions tool message, a model without image input, an unmaterialized stream source) is dropped at its own position with a reason, and the model receives an in-place [omitted tool output: ...] placeholder so the surrounding text keeps its meaning.

Streaming

GenerateStream returns a stream of deltas plus the final result. Streaming is provider-neutral; each driver adapts its native protocol.

Media streams

Live media input is a first-class transport, not a new DTO: a stream is a sequence of ordinary message.Parts. The media layer owns the generic pull contract (media.Stream[T] / media.Pipe[T]); message.Stream is that contract instantiated over Part, and message.NewPartPipe builds a bounded pipe whose Send blocks when the buffer is full — that is the backpressure contract. Interrupt aborts a stream (barge-in, error), while Close ends it normally; after Interrupt, Read returns context.Canceled even if buffered parts remain.

An AudioSource or VideoSource can carry a live stream via message.NewAudioStream / message.NewVideoStream (source kind stream). Stream sources are valid only while a message is in flight:

Transcription

Speech recognition is a first-class workload with two execution shapes:

Both shapes address the same ModelRef and share the Transcription route pools. Drivers that only serve one shape leave the other opener nil and the assembly reports UnsupportedOperation for it.

Sessions may emit multiple Final events for continuous recognition. A provider session can expose the optional TranscriptionSessionFinisher capability: callers that have no more audio call FinishInput, then drain Next to io.EOF and read Result. TranscribeStream performs that end-of-input handshake automatically after the source stream ends; the script bridge exposes it as the session handle's finish().

Live input rides the part-stream transport: FeedTranscription pumps a message.Stream[Part] into an open session (audio parts become chunks with monotonic sequence; EOF ends feeding; a stream failure interrupts the session), and TranscribeStream is the one-shot open + feed + drain + result form. Unary Transcribe rejects stream sources — whole-file recognition takes complete audio, live audio goes through a session.

See graph.md for the inference node.