Guides
Route model traffic through the gateway
Change one base URL and every model request is metered, routed under your policy, checked against budgets before the call, and recorded. Your agents’ context is observed, so the gate judges what actually entered it.
Immiscible speaks the two wire protocols that carry nearly all inference: OpenAI chat completions at /v1 and Anthropic Messages at /anthropic/v1. Anything with a configurable base URL is governed in path. Streaming, tool use, prompt caching, extended thinking and image input pass through untouched.
Immiscible never resells inference. Your traffic is served on your own provider contract, with your own key, sealed with AES-256-GCM under Settings, Connections in the console.
#Two reasons to do it
- Spend. Routing, budgets scaled by yield, attribution to people, teams and units of work, and a per-request record of what ran, what it cost and where the data went.
- Observed provenance. The gateway reads each request’s message history and records, per session, what untrusted content entered the context: web fetches and searches, email, documents and third-party MCP tool output. When the agent then asks the gate, Immiscible compares what it declares with what the gateway saw. See observed provenance.
#1. Issue a key
Under Settings, For engineers, API keys, issue a gateway key. It binds traffic to a person or a service, a team and optionally a default task class. An agent’s own agent key works here too, so its inference is governed and attributed to it, and the kill switch covers it: a stopped agent’s model calls are refused with 403 agent_stopped. Stop every agent refuses model calls on every key in the workspace, people’s included; see what a stop stops.
#2. Change the base URL
export ANTHROPIC_BASE_URL=https://immiscible.fly.dev/anthropic
export ANTHROPIC_AUTH_TOKEN=ask_...
# optional: declare the work
export ANTHROPIC_CUSTOM_HEADERS="x-immiscible-task-class: code.feature"# Settings, Models, OpenAI API key: paste the Immiscible key,
# turn on "Override OpenAI Base URL" and set it to:
https://immiscible.fly.dev/v1from openai import OpenAI
client = OpenAI(
base_url="https://immiscible.fly.dev/v1",
api_key="ask_...",
default_headers={"x-immiscible-task-class": "data.extract"},
)import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic({ baseURL: 'https://immiscible.fly.dev/anthropic', apiKey: 'ask_...' });Cline, Roo, Aider, Continue and Zed: choose the OpenAI-compatible provider with base URL https://immiscible.fly.dev/v1 (Aider: --openai-api-base https://immiscible.fly.dev/v1). Open WebUI, LibreChat and other chat clients: add an OpenAI-compatible connection with the same base URL.
#3. Send a request
curl "https://immiscible.fly.dev/v1/chat/completions" \
-H "authorization: Bearer ask_..." \
-H "content-type: application/json" \
-H "x-immiscible-task-class: support.triage" \
-d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Categorise: card declined."}]}'The response is the provider’s own, plus an immiscible field and x-immiscible-* headers. In a workspace in enforce mode, the request is routed under your policy (costs are in millionths of a US dollar):
{
"immiscible": {
"taskId": "tsk_837f559681dd",
"callId": "cal_5fab810d-1aa",
"costMicros": 14092,
"routedTo": "nvidia/nemotron-5-340b",
"requestedModel": "openai/gpt-5.5",
"savingsVsRequestedMicros": 69608,
"taskClass": "support.triage",
"taskClassSource": "header",
"enforcement": "observe",
"explanation": "Nemotron 5 340B (open weights) clears the support.triage floor (ifeval>=80) with 25% margin; ..."
}
}enforcement is the budget ladder’s step for this call: observe, nudge or downgrade. A new workspace is in shadow mode, and there the same request goes to the model it named, with what Immiscible would have done beside it (shadow mode workspace):
{
"immiscible": {
"taskId": "tsk_38c33694ef3e",
"callId": "cal_61825622-cfc",
"costMicros": 83710,
"routedTo": "openai/gpt-5.5",
"requestedModel": "openai/gpt-5.5",
"enforcement": "shadow",
"shadow": { "wouldRoute": "nvidia/nemotron-5-340b", "wouldBudgetAction": "observe", "counterfactualCost": 14092 }
}
}The headers say the same: x-immiscible-routed-to in enforce mode, x-immiscible-would-route-to in shadow mode, with x-immiscible-enforcement, x-immiscible-task-id, x-immiscible-call-id and x-immiscible-cost-micros. The call appears under Activity within a second.
#Shadow mode first
Every workspace starts in shadow mode. Each request goes to exactly the model it asked for; nothing is blocked, rerouted or refused, not even when a budget is exhausted. Immiscible records what ran and what it cost, what it would have routed to under your policy and at what price on the same token counts, and what the budget ladder would have done.
After a week, Assessment turns that into findings. When you are ready, switch to Enforce in Rules, Models and enforcement. The change is a ledger record.
#Task classes
Most clients cannot set a header, so the class of work resolves in this order, and every record says which rule applied:
- the
x-immiscible-task-classheader - the default bound to the key
- the default for the recognised client, from its User-Agent
- otherwise refused in enforce mode (
400 unclassified_request) and recorded as unattributed in shadow mode
Unclassified traffic is refused on purpose: spend that cannot be attributed to a unit of work cannot be yield-accounted. GET /v1/task-classes lists them.
#Routing and objectives
The router scores every eligible model on cost, capability, compliance and jurisdiction, for the task class and your policy profile.
- Compliance and jurisdiction are gates, not weights. A model that fails one is excluded before scoring, never traded off against price.
- Capability is a floor with margin. The cheapest model that barely clears the floor is not chosen.
- Jurisdiction is three questions: who trained the weights, who serves them (which decides CLOUD Act reach), and where the bytes land.
- Protocol affinity. A request with tools, cache markers, images or thinking blocks is routed only among models that speak its protocol.
Set an objective per request with x-immiscible-objective: balanced (the default), cost, quality, latency or sovereign. Policy profiles are pragmatic, eu-sovereign, uk-public, healthcare, permissive and startup. When nothing qualifies, the answer is 409 policy_conflict naming the binding constraint and what relaxing it would admit, at what cost, signed off by whom. POST /v1/route/preview shows the choice without calling a model.
#Budgets
A budget belongs to a team, a person or the workspace and runs for a calendar month. It is the amount you set. Self-adjusting is opt-in: turn it on for one budget (adaptive: true) or for the workspace (adaptiveBudgets in settings), and a team’s allocation then moves between 0.7 and 1.5 times the amount set with its yield relative to the rest of the organisation, never above an optional hard ceiling.
| Utilisation | What happens |
|---|---|
| under the first threshold | observe |
| approaching the allocation | nudge: an x-immiscible-advisory header |
| near the allocation | downgrade to cheaper eligible models |
| at the allocation | 402 approval_required, naming the owner |
| past the ceiling | 429 budget_exhausted |
Each request is checked against the projected position and its estimate is reserved before the call, so two concurrent requests cannot both take the last pound. Settlement uses the provider’s reported usage, never the estimate.
#Close the loop with outcomes
Accounting is by task, not request. A task is one unit of work, grouped from an x-immiscible-task-id header, the client’s session id, or calls on the same key and class less than 30 minutes apart. Report the outcome and spend becomes yield:
curl -X POST "https://immiscible.fly.dev/v1/outcomes" \
-H "authorization: Bearer ask_..." -H "content-type: application/json" \
-d '{ "taskId": "tsk_b5f8c335e4a6", "status": "accepted" }'For coding work, connect GitHub under Settings, Connections and merged pull requests report themselves: a merged PR marks linked tasks accepted, one closed without merging marks them rejected. Every delivery is verified against x-hub-signature-256.
#Sensitive data in prompts
Before a prompt leaves for a model, its text is scanned for card numbers (Luhn-checked), IBANs (mod-97), UK National Insurance and US Social Security numbers, email addresses and phone numbers. The prompt DLP setting under Agents, Vault (dlp in the API) decides what happens: off, record (the default: counted, prompt unchanged), redact (text replaced with tokens such as [card ••••4242]) or block (400 sensitive_data_blocked). Matched values are never stored in clear.
#Observed provenance
For each session, the gateway keeps the kind of untrusted source that entered the context, the tool name, a SHA-256 digest and the time. Never the content, and only a digest of the session id.
- If the agent declares only trusted sources but the session observably carries untrusted content,
provenance_mismatchfires and a person decides. - The Rule of Two is judged on declared and observed sources together.
- Taint is monotonic: a later clean request does not clear it. Sessions expire seven days after last seen.
- A request naming no session, or one the gateway never saw, is judged against everything that agent’s own key carried in the last 30 minutes, so an agent cannot shed taint by quoting a fresh session id.
The gateway reads the session id from x-immiscible-client-session, then Claude Code’s metadata.user_id, then OpenAI’s user field. A client that names none is given one in x-immiscible-session and sends it back. For your own agent, send the same id as x-immiscible-client-session on model calls and as session.id on action requests.
#What the gateway cannot see
Devin, GitHub Copilot and Cursor’s own hosted models call inference from their vendors’ backends, where no gateway can sit. Their usage is reconciled, not governed: pulled from the vendor’s API (for Devin, POST /v1/ingest/devin) and attributed to people, teams and outcomes, marked governed: false, and counted as a gap in compliance coverage rather than dressed up as coverage.