RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Prefix caching

Memory

Prefix caching

What it is

Prefix caching preserves KV cache blocks for prompt prefixes that recur across requests. A matching request can reuse the cached prefix and compute only the uncached suffix.

Why it moves cost and latency

Reuse removes repeated prefill work for shared instructions, templates, or documents. It can reduce first-token latency and compute cost when cache hits are frequent and prefixes are long enough.

What it looks like in practice

The server identifies reusable prefixes, stores their cache blocks, and checks later requests for exact matches. Gains depend on prompt repetition, cache capacity, eviction behavior, and the cost of lookup.

Where we measured it

  • Prefix-cache throughput by workload->

Related terms

  • Paged attention->
  • KV cache->
  • TTFT->
  • Concurrency->
  • Attention sinks->

Questions this definition answers

What does Prefix caching mean in inference serving?

Prefix caching preserves KV cache blocks for prompt prefixes that recur across requests. A matching request can reuse the cached prefix and compute only the uncached suffix. The server identifies reusable prefixes, stores their cache blocks, and checks later requests for exact matches. Gains depend on prompt repetition, cache capacity, eviction behavior, and the cost of lookup.

Why can Prefix caching move cost or latency?

Reuse removes repeated prefill work for shared instructions, templates, or documents. It can reduce first-token latency and compute cost when cache hits are frequent and prefixes are long enough.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy