Your AI does thousands of jobs a day. Nobody checks the work.
Metrx finds cheaper, better configurations for every AI workload, proves them against a randomized holdout, and switches on your say-so — as models and prices change.
- Pre-registeredcontract signed
- Randomized trialvs. holdout
- Causal readouteffect estimated
- Promotedbreakers armed
Watch the loop run, end to end.
37 seconds: what Metrx does, over the same surfaces a customer sees. Every figure in the film is synthetic demo data. Then open the real read-only workspace — no signup.
Every AI workload was tuned once — and then the world moved.
Configurations rot silently.
Models get deprecated, prices reshuffle, prompts age against new traffic. The config that was right at launch quietly stops being right — and nothing tells you.
Benchmarks aren't your traffic.
Public leaderboards say what a model can do — not what it does on your workloads, your prompts, your users. The only benchmark that matters is production.
Nobody dares touch prod.
Without controlled trials and a rollback story, every config change is a leap of faith. So teams either freeze — and overpay — or gamble on vibes.
From guesswork to evidence.
You find out a model got cheaper from a changelog.
A candidate is generated the week prices move, and tested before you read the changelog.
Someone swaps a model and everyone hopes.
The swap runs against a randomized holdout, judged by the contract you approved.
"Is this actually better?" gets answered with a vibe.
It gets answered with an effect estimate and a reproducible record.
A regression is discovered by a customer.
A breaker trips on the contract and the last proven config comes back.
Two loops, always running.
One loop hunts for better configurations and proves them. The other guards what is already live. Together they keep every workload on its best known footing.
Find, prove, and apply better configurations.
Metrx generates candidates for each workload — model, prompt, parameters, routing — and is built to put them through pre-registered trials on live traffic.
| Workload | Candidate | Δ cost | Δ quality | Status |
|---|---|---|---|---|
| support-triage | cfg-cand-11 | −43.9% | +0.3 pt | Synthetic |
| doc-extraction | cfg-cand-07 | −21.5% | −0.1 pt | Attributed |
| lead-scoring | cfg-cand-19 | −2.4% | −1.8 pt | Negative · kept |
| rag-answering | cfg-cand-03 | measuring… | measuring… | Measured |
Guard every configuration that is already live.
Promotion is not the end of the story. The Assure loop watches for drift, re-checks incumbents when providers ship model releases, and rolls back when a breaker trips.
Three steps to a proven configuration.
Connect a workload
Point the SDK at an agent. Metrx sees the traffic, the cost, and the shape of the work.
Approve an acceptance contract
You define what "better" means for that workload — quality floor, guardrails, latency and budget limits. Nothing promotes outside it.
Let the loops run
Candidates are proposed and, once your traffic has evidence, proven against a holdout. Winners go live under the authority you set; losers stay in the ledger.
Trust the procedure, not the pitch.
Metrx does not ask you to believe its numbers. It shows how each number was produced, and makes the procedure itself the thing you audit.
Candidate generation is separated from evaluation.
The system that proposes a configuration never grades its own work. Evaluation runs independently, against criteria fixed in advance.
You approve the objectives and the acceptance contract.
What counts as better is your call, written down before any trial starts. We optimize inside it.
Treatment is randomized; results are causal.
Candidates are designed to run against a randomized holdout on live traffic, so a measured effect is an effect — not a before/after coincidence.
Negative and inconclusive results are preserved.
Failed candidates stay in the ledger with full readouts. Never retried until positive, never quietly deleted.
Every promotion carries a reproducible evidence record.
Inputs, splits, exclusions and readouts are sealed and addressable — re-runnable by you or your auditors.
Built for how you run AI.
Start free. Move to Platform when the engine has earned it.
Cost analytics and a sample verdict cost nothing. Platform pricing is set per workspace at contract from your trailing spend.
The things people ask first.
How is this different from an eval tool?
Traces and evals tell you what happened. Metrx is built to run randomized trials on your traffic, state the result in dollars against a holdout, and put the winner live under the authority you set.
How is this different from a router?
A router picks a model from population benchmarks. Metrx is built to prove what works on your traffic, once it has evidence — and can meter whether the router’s claim survived contact with it.
Do you need my prompts and data?
Metrx works from metrics about your calls — cost, tokens, latency, and the outcome signal you choose to send. What a workload exposes is part of its acceptance contract.
What if a change makes things worse?
That is what the holdout, the exposure budget and the breakers are for. Trials run on a bounded slice of traffic, and the incumbent configuration is restored when a contract is violated.
What does it cost?
Start free. Platform pricing is set per workspace at contract from your trailing spend — see the pricing page.
Stop guessing. Head every workload toward its best proven configuration.
Connect a workload, approve an acceptance contract, and let the loops run.
Every figure shown on this page is synthetic demo data, labeled as such.