Video summary
Complete DSPy Course | Automatic and Programmatic Prompt Optimization | Complete Course
Main summary
Key takeaways
Overview
This video is a complete walkthrough of prompt engineering and prompt optimization for LLM-based applications, focusing on DSPy (referred to as “DSP” / “DSPI” in the subtitles). It shows how DSPy automates and structures prompt improvements.
Why prompt optimization matters (scalability + cost)
The speaker has scaled AI systems to millions of prompts and large-scale document/image parsing using Python, so prompt optimization is crucial because:
- Prompt quality is highly sensitive
- Small prompt differences can noticeably change accuracy.
- There’s a cost/performance tradeoff
- A better prompt may allow using a cheaper/smaller model.
- The core takeaway: prompts require systematic experimentation and evaluation rather than intuition alone.
Data + task used throughout: receipt tax extraction
The tutorial repeatedly uses the same benchmark task:
- Input: images of receipts
- Output: extract before-tax and after-tax amounts
- Some receipts may not explicitly contain one value, so the system may need to deduce/calculate it.
Example behavior described:
- If “before tax” isn’t present, compute it as:
before-tax = after-tax - taxes
Evaluation data:
- Ground truth exists for 11 receipts (used for evaluation).
Building a baseline system (manual prompt + parsing)
Step 1: Call an LLM from Python using LiteLLM
The baseline uses LiteLLM to:
- Easily swap model providers while keeping message structure.
Models mentioned:
- Llama Scout via Groq
- Fast/cheap, but limited context
- Fallback via OpenRouter
- Requires an API key
The baseline sets:
temperature = 0for deterministic outputs
Step 2: Enforce structured output
Baseline problem:
- Returning free-form text is hard to parse reliably.
Baseline improvement:
- Prompt the LLM to return numbers inside XML tags.
- Parse results using:
- Regex (extract values from XML)
- Pydantic to validate/clean numeric fields
- before/after tax modeled as floats
- fields allowed to be
Nonewhen missing
- The workflow also includes verification, since LLMs can be weak at math (e.g., checking calculations).
Evaluation pipeline with PixelTable
The video uses PixelTable to organize experiments and compute metrics.
PixelTable is described as:
- Like a database/dataframe, but designed to handle multimedia (images).
Workflow:
- Store receipt images in a table with image fields
- Add computed columns that trigger LLM calls and parsing
- Define a metric function comparing predictions to ground truth
Metric definition (concept)
- Returns 1 only if both:
- before-tax matches ground truth, and
- after-tax matches ground truth
- Otherwise returns 0
- Run across all rows to get accuracy (e.g., “2 out of 11 correct”).
Key finding shown
- Even tiny prompt edits can cause large accuracy drops.
- An example prompt variant performs very poorly until it’s corrected.
Automatic prompt optimization “by hand” (M.E.R.O-like approach)
Before using DSPy, the speaker builds a simplified optimizer to understand the mechanism.
M.E.R.O v2 idea (as described)
Two LLM components:
- Proposer
- Given context + failures, generates a new instruction/prompt
- Evaluator
- Tests the new instruction on examples and scores it
Context given to the proposer includes:
- Program awareness
- summary/understanding of the code being optimized
- Dataset summary
- describes table columns/inputs/outputs
- Failure cases
- aggregates wrong examples
- Instruction history + “tips” for the proposer LLM
Results
- Manual optimization improves accuracy over the baseline.
- The reported progression includes:
- starting from 2/11
- reaching up to 5/11 via iterative prompt proposals
- further improvements beyond that through additional adjustments.
Switch to DSPy: structured prompting via Signatures + Predict
Why DSPy is highlighted
DSPy avoids brittle “one big prompt string” systems by using:
- Signatures to define inputs/outputs
- DSPy-generated prompt structure via adapters/templates
- Built-in parsing/validation
- Tools for retrying/templating alternatives
Signature-based program
The speaker defines a DSPy “signature” conceptually like:
- Input: receipt image
- Output: receipt totals (before/after tax) as a structured type
Then wraps it in a DSPy program (e.g., Predict(...)) connected to a chosen LLM.
Baseline DSPy behavior
- DSPy without optimization still resembles baseline performance (example given: 2/11).
DSPy automatic prompt optimization (M.E.R.O inside DSPy)
Training set + metric + optimizer
DSPy requires an example object format, including:
receipt_image(DSPy image object)receipt_totals(ground truth)
Then:
- Define a DSPy metric mirroring the earlier metric
- Configure the DSPy optimizer to use:
- M.E.R.O style instruction optimization
- a teacher LLM (e.g., Gemini) to propose improved instructions
- a task/program LLM (e.g., Llama Scout) to perform predictions
Finally:
- Run compilation to produce an optimized DSPy program/prompt.
Outcomes
- DSPy automatic optimization improves accuracy (reported around 5/11, with later gains after subsequent work).
Few-shot optimization attempt
DSPy is also used to try few-shot prompting (adding training demonstrations).
The video reports mixed/negative outcomes:
- Few-shot was more expensive
- large context usage, image-heavy inputs
- Sometimes it did not improve accuracy beyond instruction-only optimization
Final takeaway:
- Few-shot can be tested quickly, but is often not worth the cost for this case.
Manual prompt engineering “at the program level” (DSPy modules)
After automatic optimization, the speaker explores DSPy “knobs,” including:
- Reasoning mode
- chain-of-thought / structured reasoning templates
- React/agent behavior
- tool calling loops
- Template/adapters
- affect prompt formatting (e.g., chat adapter vs JSON adapter)
Tested upgrades and results
- Reasoning added
- Improved accuracy (reported up to ~8/11 at one point)
- Agent + calculator tools
- Didn’t help; accuracy worsened
- Tool overhead/errors increased cost (multi-call loops)
- Model switching (Gemini vs Llama)
- Gemini improved further, but cost increased
Final “best” result achieved
The tutorial ends at 11/11 correct using Gemini with a specific strategy that included:
- encouraging the model to output a list of money-related numericals
- reasoning that correctly handles missing fields (including correct
Nonehandling)
Additional notes:
- Llama Scout could do well but not always to 11/11 under the same strategy (example: 8/11 shown for Llama under some settings).
- Maverick (via OpenRouter/Groq context considerations) performed worse than Scout in that experiment.
Overall conclusions emphasized
- Prompt engineering isn’t dead—DSPy makes it systematic and scalable.
- DSPy is presented as a state-of-the-art workflow because it provides:
- structured prompt composition (Signatures + adapters/templates)
- automatic instruction optimization
- evaluation/iteration loops
- program saving/loading for reuse
- control over formatting and behavior (reasoning, agents, tool calling)
- Key operational message: use evaluation-driven iteration and test changes empirically rather than guessing.
Named main speakers / sources
- Main speaker: the tutorial author (not explicitly named in subtitles)
Tool/library references used
- DSPy (DSP / “Aspire” mentioned as DSP equivalents)
- M.E.R.O / M.E.R.O v2 (inspiration/optimizer approach)
- LiteLLM
- Pydantic
- PixelTable
- OpenAI (mentioned as prompt engineering reference/documentation)
- Anthropic
Model providers mentioned
- Groq
- OpenRouter
- Examples: Gemini, Llama Scout, and Maverick