Skip to main content
A benchmark report measures your optimized winner and its baseline through the same pinned harness, on named public benchmarks, and returns both sets of numbers together. It is the answer to “the run says it is faster, but is it still as good”. This is real GPU work on real evaluation sets, so it is priced and confirmed before it starts.

The Benchmarks tab

The Benchmarks tab in your session is the supported way to start a report. It is scoped to one optimized version of one pipeline, and it offers the benchmark suites wired for that model’s modality. Select the version, choose the suites, and use Run benchmarks. The tab shows live progress while the run is in flight and the finished report when it settles. Two conditions gate the button:
  • An optimization is still running. Benchmarks unlock when it completes. A report measures a settled winner, so there is nothing coherent to measure mid-run.
  • A benchmark is already running on this pipeline. One live benchmark run per pipeline. You get the running one’s status rather than a second one.

Asking for a benchmark in chat

You can ask for a benchmark in the session chat, and the agent will handle it. What it does not do is start it on its own. Instead it prices the report through the same path the Benchmarks tab prices through, and hands you a confirmation carrying two figures:
  • the expected cost of the run, and
  • the amount held until the run settles.
Both are shown, always together. The larger figure is a hold, not an estimate, and it is never presented as one. Your confirmation in the session is what authorizes the spend and starts the run. Until you confirm, nothing has started and nothing has been charged.
Pricing a report is free and leaves no record. The refusals that apply to a report apply at pricing time too, so a benchmark that cannot run is refused before you are ever shown a price for it, rather than after you pay for it.

Reading a report from chat

Chat can also read, not just price. Ask about a benchmark that already ran and you get the finished report replayed from the stored record, which is the same payload the Benchmarks tab renders. Ask about one still in flight and you get its real phase counts. Progress is reported as counts of benchmark phases completed and running. RunInfra does not publish a percentage or an estimated finish time for a benchmark run, because neither can be derived honestly from phase counts alone.

Starting one from chat is refused

A benchmark cannot be dispatched directly from a chat message, and an attempt to do so is refused by name before any benchmark work begins. Nothing is charged. The reason is a deliberate separation. The billing authorization that a paid benchmark needs is minted by your confirmation in the session, so a path that skips the confirmation is a path with no authorization attached to it. Rather than run work that could not be billed to anyone, RunInfra refuses and tells you the supported route. Chat prices and reads. The session confirmation starts. From that point the chat path and the tab path are the same run, writing to the same record.

While a run is live

A benchmark report is blocked while a runbook execution is live on the same pipeline, and the refusal says so:
A runbook execution is live on this pipeline right now, so the benchmark report is blocked until it settles. Nothing was started.
The block is temporary and self-clearing. Once the current run finishes or settles, ask again and the report runs. Being blocked by a live run never cancels, pauses, or otherwise disturbs that run.