Skip to main content
A benchmark report measures your optimized winner and its baseline the same way, on the same pinned instrument and the same named public benchmarks, and returns both sets of numbers together. It answers the question a run leaves open: it is faster, but is it still as good. This is real measurement work on real evaluation sets, so it is priced and confirmed before it starts.

Run one

1

Open the Benchmarks tab

It is scoped to one optimized version of one pipeline, and it offers the benchmark suites wired for that model’s modality.
2

Select the version and the suites

Then use Run benchmarks. You see the price and your available balance before anything starts.
3

Watch it, then read it

The tab shows live progress while the run is in flight, and the finished report when it settles.
Two things gate the button, and both clear on their own:
  • An optimization is still running. A report measures a settled winner, so there is nothing coherent to measure mid-run. Benchmarks unlock when it completes.
  • A benchmark is already running on this pipeline. One live benchmark run per pipeline. You get the running one’s status instead of a second run.

What it costs

One balanceone balance per workspaceBalanceCreditsone balance per workspace$1 free,once per accountAgent chat and plansCharged after each turn for the work that turn actually did.Optimization and benchmarking runsQuoted before the run, held while it runs, settled when it finishes.Your own deployed endpointsPer million tokens, input and output separately.Model APIsPer million tokens at the price published on that model’s page.Zero balance402 Payment Requiredbefore any work or charge.
One balanceone balance per workspaceBalanceCreditsone balance per workspace$1 free,once per accountAgent chat and plansCharged after each turn for the work that turn actually did.Optimization and benchmarking runsQuoted before the run, held while it runs, settled when it finishes.Your own deployed endpointsPer million tokens, input and output separately.Model APIsPer million tokens at the price published on that model’s page.Zero balance402 Payment Requiredbefore any work or charge.
Your confirmation authorizes the work. RunInfra checks that the workspace is funded, then bills actual eligible usage after the run finishes. Until you confirm, nothing has started and nothing has been charged. At confirmation, RunInfra rechecks the workspace payment hold and the benchmark’s reserved coverage before dispatch, and the same checks apply if a confirmation is retried. A payment hold blocks the run even when credits remain, and a billing check that cannot be completed refuses without starting benchmark work.

Chat prices and reads, the session starts

You can ask for a benchmark in the session chat. The agent will price it through the same path the Benchmarks tab uses, and hand you a confirmation carrying the expected cost and your available balance. What chat will not do is start it. A benchmark cannot be dispatched from a chat message, and an attempt is refused by name before any benchmark work begins. Nothing is charged. The reason is a deliberate separation: your session confirmation is what authorizes paid work, so a path that skips the confirmation has no approval to spend. From the confirmation onward, the chat path and the tab path are the same run writing to the same record. Chat also reads. Ask about a report that already ran and you get it replayed from the stored record, the same payload the tab renders. Ask about one still in flight and you get its real phase counts. RunInfra does not publish a percentage or an estimated finish time for a benchmark run, because neither can be derived honestly from phase counts.

While a run is live

A benchmark report is blocked while a runbook execution is live on the same pipeline, and the refusal says so:
A runbook execution is live on this pipeline right now, so the benchmark report is blocked until it settles. Nothing was started.
The block is temporary and self-clearing. Ask again once the current run settles. Being blocked never cancels, pauses, or otherwise disturbs that run.

Run outcomes

The quality verdict a run records on its own.

Optimization runs

How a winner gets settled in the first place.

Deployments

Serve the winner, or take the kit and run it yourself.