Video summary

What is Prompt Caching? Optimize LLM Latency with AI Transformers

Main summary

Key takeaways

Technology

Technological concept: Prompt caching (not output caching)

  • Prompt caching speeds up and reduces cost for LLM requests by caching only the input prompt’s computed internal state, so the model doesn’t need to re-process the same prompt context again.
  • It is distinct from output caching:
    • Output caching would reuse a previously generated response for an identical prompt (analogous to caching SQL query results).
    • Prompt caching instead reuses precomputed “KV pairs” (key-value attention representations) derived from the prompt, not the final output.

How it works in transformer models

  • When an LLM processes a prompt, it computes KV (key-value) pairs:
    • across every transformer layer
    • for every input token
  • This computation happens in the “prefilled” phase before the model can generate its first output token.
  • Prompt caching stores those precomputed KV pairs, so later requests can skip the expensive prefill computation for the reused prompt prefix.

Main use case examples

  • Large repeated context
    • If the prompt includes something huge (e.g., a 50-page document) and you ask it to do something like summarization, prefill becomes extremely expensive (thousands of tokens, many layers).
    • A later request can reuse the same document prefix but change only the question, saving latency and cost.
  • Most common cached content: system prompt
    • Chatbots typically send consistent system instructions (personality, rules, behavior).
    • These can be cached so each new conversation doesn’t redo the system-level prompt processing.
  • Other cacheable prompt parts mentioned:
    • Few-shot examples
    • Tool/function definitions
    • Conversation history

When caching applies: prefix matching

  • The caching system uses prefix matching:
    • It compares the new prompt token-by-token from the start against what’s already cached.
    • Caching continues until the first differing token, then normal processing resumes.
  • Prompt structure matters:
    • Static content should be placed at the beginning (e.g., system instructions → document → few-shot examples → then the user question).
    • If dynamic content (like the user question) appears first, caching may fail immediately because the prefix changes.

Practical notes / thresholds

  • Minimum size to benefit
    • Caching typically needs at least ~1024 tokens to overcome overhead.
  • Cache expiration
    • Usually cleared after ~5–10 minutes (some may persist up to ~24 hours).
  • Provider support modes
    • Some providers support automatic prompt caching (via prefix matching).
    • Others require explicitly marking which prompt parts to cache in the API call.

Main speakers / sources

  • Single narrator/speaker discussing prompt caching conceptually (no other speakers or named sources mentioned in the subtitles).

Original video