Video summary
What is Prompt Caching? Optimize LLM Latency with AI Transformers
Main summary
Key takeaways
Technological concept: Prompt caching (not output caching)
- Prompt caching speeds up and reduces cost for LLM requests by caching only the input prompt’s computed internal state, so the model doesn’t need to re-process the same prompt context again.
- It is distinct from output caching:
- Output caching would reuse a previously generated response for an identical prompt (analogous to caching SQL query results).
- Prompt caching instead reuses precomputed “KV pairs” (key-value attention representations) derived from the prompt, not the final output.
How it works in transformer models
- When an LLM processes a prompt, it computes KV (key-value) pairs:
- across every transformer layer
- for every input token
- This computation happens in the “prefilled” phase before the model can generate its first output token.
- Prompt caching stores those precomputed KV pairs, so later requests can skip the expensive prefill computation for the reused prompt prefix.
Main use case examples
- Large repeated context
- If the prompt includes something huge (e.g., a 50-page document) and you ask it to do something like summarization, prefill becomes extremely expensive (thousands of tokens, many layers).
- A later request can reuse the same document prefix but change only the question, saving latency and cost.
- Most common cached content: system prompt
- Chatbots typically send consistent system instructions (personality, rules, behavior).
- These can be cached so each new conversation doesn’t redo the system-level prompt processing.
- Other cacheable prompt parts mentioned:
- Few-shot examples
- Tool/function definitions
- Conversation history
When caching applies: prefix matching
- The caching system uses prefix matching:
- It compares the new prompt token-by-token from the start against what’s already cached.
- Caching continues until the first differing token, then normal processing resumes.
- Prompt structure matters:
- Static content should be placed at the beginning (e.g., system instructions → document → few-shot examples → then the user question).
- If dynamic content (like the user question) appears first, caching may fail immediately because the prefix changes.
Practical notes / thresholds
- Minimum size to benefit
- Caching typically needs at least ~1024 tokens to overcome overhead.
- Cache expiration
- Usually cleared after ~5–10 minutes (some may persist up to ~24 hours).
- Provider support modes
- Some providers support automatic prompt caching (via prefix matching).
- Others require explicitly marking which prompt parts to cache in the API call.
Main speakers / sources
- Single narrator/speaker discussing prompt caching conceptually (no other speakers or named sources mentioned in the subtitles).