Optimization engine (founding pilot)

Your AI does thousands of jobs a day. Nobody checks the work.

Metrx finds cheaper, better configurations for every AI workload, proves them against a randomized holdout, and switches on your say-so — as models and prices change.

Every number carries its provenance
MeasuredAttributedVerifiedSynthetic
workload / support-triage · trial TR-0847Synthetic demo data
cfg-baseline-04incumbent · frontier-large · temp 0.2
Cost / 1k tasks
$4.12
Task success
96.1%
p95 latency
2.9s
Incumbent
cfg-cand-11candidate · frontier-small + reranker · temp 0.1
Cost / 1k tasks
$2.31
Task success
96.4%
p95 latency
1.7s
Promoted
Randomized traffic splitbounded exposure
candidate 20%incumbent 70%holdout 10%
Evidence chainrecord ev_9f2c…a41d · reproducible
  1. Pre-registeredcontract signed
  2. Randomized trialvs. holdout
  3. Causal readouteffect estimated
  4. Promotedbreakers armed
Synthetic demo data

Watch the loop run, end to end.

37 seconds: what Metrx does, over the same surfaces a customer sees. Every figure in the film is synthetic demo data. Then open the real read-only workspace — no signup.

The problem

Every AI workload was tuned once — and then the world moved.

01 / DRIFT

Configurations rot silently.

Models get deprecated, prices reshuffle, prompts age against new traffic. The config that was right at launch quietly stops being right — and nothing tells you.

02 / EVIDENCE

Benchmarks aren't your traffic.

Public leaderboards say what a model can do — not what it does on your workloads, your prompts, your users. The only benchmark that matters is production.

03 / RISK

Nobody dares touch prod.

Without controlled trials and a rollback story, every config change is a leap of faith. So teams either freeze — and overpay — or gamble on vibes.

What changes

From guesswork to evidence.

Today

You find out a model got cheaper from a changelog.

With Metrx

A candidate is generated the week prices move, and tested before you read the changelog.

Today

Someone swaps a model and everyone hopes.

With Metrx

The swap runs against a randomized holdout, judged by the contract you approved.

Today

"Is this actually better?" gets answered with a vibe.

With Metrx

It gets answered with an effect estimate and a reproducible record.

Today

A regression is discovered by a customer.

With Metrx

A breaker trips on the contract and the last proven config comes back.

The engine

Two loops, always running.

One loop hunts for better configurations and proves them. The other guards what is already live. Together they keep every workload on its best known footing.

Loop 01 — Optimize

Find, prove, and apply better configurations.

Metrx generates candidates for each workload — model, prompt, parameters, routing — and is built to put them through pre-registered trials on live traffic.

workspace / active trialsSynthetic demo data
WorkloadCandidateΔ costΔ qualityStatus
support-triagecfg-cand-11−43.9%+0.3 ptSynthetic
doc-extractioncfg-cand-07−21.5%−0.1 ptAttributed
lead-scoringcfg-cand-19−2.4%−1.8 ptNegative · kept
rag-answeringcfg-cand-03measuring…measuring…Measured
Loop 02 — Assure

Guard every configuration that is already live.

Promotion is not the end of the story. The Assure loop watches for drift, re-checks incumbents when providers ship model releases, and rolls back when a breaker trips.

workspace / drift monitorSynthetic demo data
support-triagecfg-cand-11 · live 18d
In envelope
doc-extractioncfg-base-02 · live 64d
In envelope
rag-answeringcfg-base-09 · rolled back
Breaker tripped
lead-scoringcfg-base-05 · release check queued
Re-verifying
How it works

Three steps to a proven configuration.

1

Connect a workload

Point the SDK at an agent. Metrx sees the traffic, the cost, and the shape of the work.

2

Approve an acceptance contract

You define what "better" means for that workload — quality floor, guardrails, latency and budget limits. Nothing promotes outside it.

3

Let the loops run

Candidates are proposed and, once your traffic has evidence, proven against a holdout. Winners go live under the authority you set; losers stay in the ledger.

Procedural integrity

Trust the procedure, not the pitch.

Metrx does not ask you to believe its numbers. It shows how each number was produced, and makes the procedure itself the thing you audit.

P-01
Separation

Candidate generation is separated from evaluation.

The system that proposes a configuration never grades its own work. Evaluation runs independently, against criteria fixed in advance.

P-02
Contract

You approve the objectives and the acceptance contract.

What counts as better is your call, written down before any trial starts. We optimize inside it.

P-03
Randomization

Treatment is randomized; results are causal.

Candidates are designed to run against a randomized holdout on live traffic, so a measured effect is an effect — not a before/after coincidence.

P-04
Negatives

Negative and inconclusive results are preserved.

Failed candidates stay in the ledger with full readouts. Never retried until positive, never quietly deleted.

P-05
Evidence

Every promotion carries a reproducible evidence record.

Inputs, splits, exclusions and readouts are sealed and addressable — re-runnable by you or your auditors.

Where you fit

Built for how you run AI.

Start free. Move to Platform when the engine has earned it.

Cost analytics and a sample verdict cost nothing. Platform pricing is set per workspace at contract from your trailing spend.

Questions

The things people ask first.

How is this different from an eval tool?

Traces and evals tell you what happened. Metrx is built to run randomized trials on your traffic, state the result in dollars against a holdout, and put the winner live under the authority you set.

How is this different from a router?

A router picks a model from population benchmarks. Metrx is built to prove what works on your traffic, once it has evidence — and can meter whether the router’s claim survived contact with it.

Do you need my prompts and data?

Metrx works from metrics about your calls — cost, tokens, latency, and the outcome signal you choose to send. What a workload exposes is part of its acceptance contract.

What if a change makes things worse?

That is what the holdout, the exposure budget and the breakers are for. Trials run on a bounded slice of traffic, and the incumbent configuration is restored when a contract is violated.

What does it cost?

Start free. Platform pricing is set per workspace at contract from your trailing spend — see the pricing page.

Stop guessing. Head every workload toward its best proven configuration.

Connect a workload, approve an acceptance contract, and let the loops run.

Every figure shown on this page is synthetic demo data, labeled as such.