Glossary
Definitions for the serving mechanisms and metrics that change inference cost, latency, memory pressure, and quality.
Use a workspace API key and pay for input, cached input, and output tokens.