Product

Everything Runback does, in the order you would use it.

Four stages, and the workspace around them — the same grouping the product uses internally, so what you read here is what you land in.

01 · Observe

Every decision, captured at the boundary.

Every model call — context, tools, tokens — recorded at the grain a decision is actually made. PII is redacted in-process, before anything leaves your app.

The Runs list in Runback, showing five agent runs with status, step count, tokens and start time, above a fleet determinism check.

Runs · Alerts · Incidents · Coverage · Topologywhat each one does →

02 · Replay

Re-run the real decision, not an approximation of it.

Re-execute from the exact captured context — tools, retrieval and the messages the model saw all held fixed. A different output means behaviour changed, and you know which change caused it.

The Evals screen in Runback, listing eval runs with their status and scores.

Prompts · Playground · Datasets · Evals · Golden corpus · Compare · Corpus signalswhat each one does →

03 · Gate

Block the violating action before it runs.

Write a rule and simulate it against ninety days of real decisions before it is live — so you know exactly what it would have stopped, and what it would have cost you. Then enforce it, and the block is sealed into the record.

The Policies screen in Runback, listing seven active policies with versions and rule counts, above a coverage-gaps panel naming tools no rule currently guards.

Policies · Policy library · Causal attribution · Models · Model diff · Upgrade gate · Approvalswhat each one does →

04 · Audit

Produce the record on demand. Let anyone check it.

A SHA-256 hash-chained, Ed25519-signed record of every decision. Tamper-evident, and verifiable by a regulator or an auditor without an account and without installing Runback.

The Audit ledger in Runback, showing the tamper-evident seal over the run history.

Audit ledger · Compliance artifacts · Regulatory mapping · Administrative audit log · Open verificationwhat each one does →

The difference

Five things the observability tools do not do.

Runback's behavioural drift screen: a stability score of zero against a 30-day baseline, with agents monitored, unacknowledged alerts and alerts over the last 30 days.

Behavioural drift

Catch a model quietly changing upstream, with no code change on your side.

Runback's bisect screen: a run selected for search above ten candidate models ordered good to bad, from gpt-oss-120b through to claude-haiku.

Bisect

Binary-search a candidate list for the exact change that broke it — O(log n) probes, each one recorded.

Runback's judge calibration screen: a rubric with one verdict reviewed, 100% agreement and two pending, each verdict offering agree or correct.

Judge calibration

Correct the LLM judge; every correction is fed back into future judging.

Runback's fleet benchmark screen: a fintech vertical score of 63, with policy compliance, error rate, response latency and token efficiency each shown against the fintech median.

Fleet benchmarks

Your agents against anonymised peers in your vertical.

Runback's cost attribution screen: total cost, tokens and runs for the last 30 days, broken down by model, with every figure traced to a captured run.

Cost attribution

Where the spend goes, by model and agent, with per-team budget caps.

Live demo · sample data, not customer figures

Want the mechanism rather than the surface? See how it actually works →

See it against your own agents.