
Behavioural drift
Catch a model quietly changing upstream, with no code change on your side.
Four stages, and the workspace around them — the same grouping the product uses internally, so what you read here is what you land in.
Every model call — context, tools, tokens — recorded at the grain a decision is actually made. PII is redacted in-process, before anything leaves your app.

Runs · Alerts · Incidents · Coverage · Topology — what each one does →
Re-execute from the exact captured context — tools, retrieval and the messages the model saw all held fixed. A different output means behaviour changed, and you know which change caused it.

Prompts · Playground · Datasets · Evals · Golden corpus · Compare · Corpus signals — what each one does →
Write a rule and simulate it against ninety days of real decisions before it is live — so you know exactly what it would have stopped, and what it would have cost you. Then enforce it, and the block is sealed into the record.

Policies · Policy library · Causal attribution · Models · Model diff · Upgrade gate · Approvals — what each one does →
A SHA-256 hash-chained, Ed25519-signed record of every decision. Tamper-evident, and verifiable by a regulator or an auditor without an account and without installing Runback.

Audit ledger · Compliance artifacts · Regulatory mapping · Administrative audit log · Open verification — what each one does →

Catch a model quietly changing upstream, with no code change on your side.

Binary-search a candidate list for the exact change that broke it — O(log n) probes, each one recorded.

Correct the LLM judge; every correction is fed back into future judging.

Your agents against anonymised peers in your vertical.

Where the spend goes, by model and agent, with per-team budget caps.
Live demo · sample data, not customer figures
Want the mechanism rather than the surface? See how it actually works →