Skip to content

Answers

How do I control LLM costs per team?

Send model traffic through a gateway that checks a budget before each call. Immiscible’s gateway takes OpenAI-shaped and Anthropic-shaped requests on your own provider keys, and holds monthly budgets per team, person or workspace that nudge, downgrade, ask the owner and then stop.

Send model traffic through a gateway that checks a budget before each call, and give each team its own budget. Immiscible’s gateway takes OpenAI-shaped and Anthropic-shaped requests on your own provider keys, records cost per request, task, person and team, and holds monthly budgets that nudge near the limit, route to cheaper eligible models, ask the budget’s owner at the allocation, and refuse past a hard ceiling.

#How do I set it up?

  1. Connect the providers: Settings, Connections, paste your OpenAI and Anthropic keys. Traffic runs on your own contracts; Immiscible never resells inference.
  2. Change one base URL in each client, using an Immiscible key instead of the provider’s:
Shell
export OPENAI_BASE_URL=https://immiscible.fly.dev/v1             # the OpenAI SDK and compatible tools
export ANTHROPIC_BASE_URL=https://immiscible.fly.dev/anthropic   # Claude Code and the Anthropic SDK
  1. Set a budget for each team. In the console, or with an admin-scoped gateway key. Amounts are millionths of a US dollar, so 500000000 is $500:
Shell
curl -X POST "https://immiscible.fly.dev/v1/admin/budgets" \
  -H "authorization: Bearer $IMMISCIBLE_ADMIN_KEY" -H "content-type: application/json" \
  -d '{ "scope": "team", "scopeId": "'"$TEAM_ID"'", "baseAllocation": 500000000, "hardCeiling": 600000000, "ownerId": "platform-lead@example.com" }'

scope is team, principal (one person) or org (the workspace); scopeId names the team or person, and ownerId is who is asked at the allocation.

  1. Run in shadow mode first. Every workspace starts there: nothing is blocked, not even an exhausted budget. After a week, read Assessment, then switch to Enforce. See shadow mode first.

#What happens when a team reaches its budget?

Where the team isWhat the gateway does
approaching the allocationadds an x-immiscible-advisory header
near the allocationroutes to cheaper eligible models
at the allocation402 approval_required, naming the budget’s owner
past the hard ceiling429 budget_exhausted

Each request’s estimate is reserved before the call, so two concurrent requests cannot both take the last pound, and settlement uses the provider’s reported usage. See budgets.

#Can I set a budget per user or per API key?

Per person, yes: a principal budget. Budgets attach to a team, a person or the workspace rather than to a key. Your provider’s own project limits are worth keeping as well, as a backstop at the provider.

#How do I find spend that does not go through the gateway?

Discovery reads the OpenAI, Anthropic and OpenRouter admin APIs for keys and projects used outside the gateway, with their owners and spend, and brings each under control in one click or proposes revoking it (two owners). Finance dashboards put spend by team, provider and agent where finance already looks.

#What does it not do?

  • Inference that runs in a vendor’s own backend (GitHub Copilot, Cursor’s hosted models, Devin) cannot pass through any gateway; it is reconciled from the vendor’s API and marked as not governed. See what the gateway cannot see.
  • It is not a model host or reseller, and it does not score answer quality.
  • If you only need routing, caching and cost reports, a dedicated LLM gateway may be all you need; see compare. Immiscible can also route through OpenRouter.