Security
Threat model
What Immiscible defends against, the rules it enforces in code, and the limits we name before someone else does.
The threat that matters is not a rogue model. It is an ordinary, well-behaved agent reading something written by someone else and doing what it says. That is prompt injection, and it is not solved: in October 2025 researchers from OpenAI, Anthropic and Google DeepMind tested twelve published defences with adaptive attacks and bypassed every one, most with success rates above 90% (summary).
The incidents so far follow one pattern. In EchoLeak (CVE-2025-32711) a single crafted email caused Microsoft 365 Copilot to send internal data out through a trusted domain, with no click. In Comet, an agentic browser asked to summarise a page followed instructions hidden in it, including reading the user’s email from another tab (Brave). In both, the agent held three things at once: untrusted input, access to something valuable, and a way to act on the world.
Immiscible is built on removing that combination, and on making sure that when it cannot be removed, a person decides.
#Rules enforced in code
No standing credentials. An agent never holds a card number, a password, a tool’s token or a personal data field it has not just been granted. It holds an agent key and asks each time. A leaked agent key can ask; it cannot take. The MCP proxy and card rail make this structural.
The Rule of Two. Following Meta’s Agents Rule of Two, a session should combine at most two of: untrusted input, access to sensitive systems or data, and the ability to change state or communicate externally. When a request would combine all three, a person decides, whatever the mandate says.
Provenance defaults to untrusted, and is observed where possible. A request with no provenance is treated as if a stranger wrote it. With the gateway in path, what entered the agent’s context is recorded, and a declaration that contradicts it is a question or a refusal.
Fail closed. If Immiscible cannot decide, the answer is deny. The card rail declines; the hook refuses; the proxy sends nothing.
Machines stop, people start. Service tokens and identity provider signals can freeze agents, under ceilings; only people lift holds, and nobody lifts one on an agent that acts for them.
Separation of duties. Nobody approves above the line for an agent that acts for them, promotes their own agent, or confirms their own change to a card issuer, an application’s origin or a tier override.
Identity that cannot be rewritten. Agent, application and single sign-on identifiers are immutable at the database. Every record names its subject and trace.
Evidence that shows a rewrite. Chained, checkpointed, signed; verifiable offline with keys you kept. See evidence.
#Assets and adversaries
| Adversary | Wants | Stopped by |
|---|---|---|
| Content an agent reads (web page, email, document, tool output) | to steer the agent into paying, leaking or acting | Rule of Two, observed provenance, injection language, lookalike domains, recipient allowlists |
| A compromised or careless agent | to act beyond its authority | mandates, tiers, entitlements, the proxy holding credentials, the card needing a receipt |
| A leaked agent key | to take | it can only ask; freeze in one tap; inference stops too |
| A compromised service token | to stop everything, or start something | freeze ceilings and pending freezes; tokens cannot lift, approve or write custom mandates |
| An insider | to approve their own agent’s request, or widen access quietly | separation of duties, dual control, signed and chained records, reviews and recertification |
| Whoever runs the database | to rewrite what happened | hash chain, checkpoints you kept, keys you kept |
#The limits
Name them before someone else does.
- Declared provenance is the agent’s word, unless the gateway saw the session. An agent whose model traffic does not go through Immiscible, and which labels a web page as its user’s instruction, defeats the Rule of Two for that request. Mandate limits still hold, and content checks still run.
- It only sees what is routed through it. An agent holding real card details, or acting through a connector Immiscible is not part of, is not governed. Removing standing credentials is what closes that gap, and it is yours to apply.
- Hosted agents are reconciled, not governed. Inference inside a vendor’s own backend cannot be put in path.
- A decision is a control, not a guarantee. It does not move money, does not stand behind a merchant, and does not replace your bank’s fraud protection or chargeback rights.
- Pattern detection finds data with a shape. It does not find a diagnosis written in a sentence, or a customer’s name.
- One writer, one host. The SQLite deployment does not fail over across hosts by itself. See deploy and backups.
#Reporting a vulnerability
See https://immiscible.fly.dev/.well-known/security.txt. We publish what red-team rounds found and fixed in the changelog.