RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Speculative decoding

Latency

Speculative decoding

What it is

Speculative decoding uses a cheaper draft process to propose several candidate tokens before the target model verifies them. Exact acceptance methods preserve the target model's output distribution while accepting valid candidates in groups.

Why it moves cost and latency

Verification can replace multiple sequential decode steps with fewer target-model passes when draft acceptance is high. Draft overhead and rejected candidates can remove the benefit on mismatched workloads.

What it looks like in practice

A serving configuration pairs a target model with a compatible draft method and tracks acceptance alongside latency. Teams test real prompts because sequence behavior and hardware balance determine whether speculation pays off.

Related terms

  • KV cache->
  • Inter-token latency->
  • Tokens per second->

Questions this definition answers

What does Speculative decoding mean in inference serving?

Speculative decoding uses a cheaper draft process to propose several candidate tokens before the target model verifies them. Exact acceptance methods preserve the target model's output distribution while accepting valid candidates in groups. A serving configuration pairs a target model with a compatible draft method and tracks acceptance alongside latency. Teams test real prompts because sequence behavior and hardware balance determine whether speculation pays off.

Why can Speculative decoding move cost or latency?

Verification can replace multiple sequential decode steps with fewer target-model passes when draft acceptance is high. Draft overhead and rejected candidates can remove the benefit on mismatched workloads.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy