Video summary
Evaluating fine-tuned LLM using Ollama
Main summary
Key takeaways
Summary (tech concepts, product features, evaluation/guide details)
This lecture continues an instruction fine-tuning project for a small LLM and then focuses on evaluating instruction-tuned models.
1) Fine-tuning recap (stages 1–2)
Stage 1: Data preparation
- Uses an instruction dataset of ~1100 instruction input/output pairs.
- Builds training utilities:
- downloading
- batching
- creating data loaders
- Uses token-ID batching logic so batches have consistent token lengths.
- Dataset split:
- 85% train
- 10% test
- 5% validation
Stage 2: Fine-tuning
- Loads a pretrained GPT-2 checkpoint: 355M parameters (often referred to as “gpt2 medium”).
- Fine-tunes on the instruction dataset so the model learns improved instruction-following.
- Shows training/validation loss behavior.
- Due to CPU compute limits, runs for 1 epoch
- Suggestion: 2+ epochs (with better hardware) typically improves results.
2) Why evaluation is needed (stage 3: evaluation)
Even after fine-tuning, the model can produce semantically wrong but plausible responses.
Example failure (active → passive):
- Model: “prepared by the chef…”
- Ground truth: “cooked by the chef…”
The lecture emphasizes that evaluation isn’t just a simple yes/no label task. Instead, you often need to:
- compare generated text vs. ground truth
- apply a scoring system for response quality
3) Step 6: Extract and save model responses for the test set
After fine-tuning, the code:
- runs the model on the entire test subset
- decodes token IDs back to text
- removes the repeated prompt/instruction portion
- collects outputs for later analysis
It also demonstrates qualitative checks on 3 test samples, showing:
- sometimes outputs are close/mostly correct
- sometimes outputs are off-topic or incorrect
- many errors are attributed to training for only 1 epoch
4) Create an evaluation-ready dataset file
The lecture builds a new JSON file: instruction_data_with_response.json, including:
instructioninputtrue outputmodel response
This format supports later automated scoring across many items.
5) Save the fine-tuned model
To avoid re-training, checkpoint saving is highlighted:
- save with
torch.save(model_state_dict) - later restore with
load_state_dict
Example saved filename resembles: gpt2-medium-355M...pth.
6) Step 7: Quantitative evaluation using another LLM + Ollama/AMA
The lecture outlines three evaluation approaches for instruction-tuned LLMs:
- Benchmarking general knowledge using MLU (Massive Multitask Language Understanding)
- Example: MMLU uses 57 tasks across domains (STEM, humanities, etc.).
- Human preference comparison
- humans rate competing model outputs
- LLM-as-a-judge
- a strong instruction model scores responses vs. ground truth
- the lecture implements this approach (#3)
How scoring is automated (LLM-as-judge)
- Uses Llama 3 Instruct (8B) as the judge model.
- Runs locally via:
- Ollama (model hosting/inference on a laptop)
- “AMA” (mentioned as an inference application/interface; text generation only)
Install/run examples:
- macOS/Linux:
ollama run llama3 - Windows alternative:
ollama servethenollama run llama3
Python helper concept:
query_model(prompt, model)to send prompts to Ollama-served Llama 3 and retrieve scores.
Scoring prompt behavior
- The judge is prompted with:
- instruction/input
- ground-truth output
- model response
- The judge returns a score, typically requested on a 0–100 scale.
- The lecture demonstrates:
- first requesting rationale
- then modifying the prompt to request integer-only scores
Compute constraints
- CPU evaluation is slow and memory-heavy:
- prompting many examples can cause laptop hangs
- full test-set evaluation may be impractical
- Suggestion: use GPU or higher RAM.
7) Results and interpretation
For 3 demo items scored by the judge:
- One response gets a high partial score (~85) despite style differences
- One response gets low (~20) because it doesn’t answer correctly (e.g., thunderstorm cloud type)
- One response gets zero because the output is factually wrong (e.g., Pride and Prejudice author)
Notes:
- Scores may not be perfectly deterministic (minor variation possible).
8) Reported average performance + improvement directions
- Average score is ~just above 50
- context: limited training data, 1 epoch, CPU constraints
- Observes common issues:
- incorrect copying of output text
- failure to follow the intended instruction
Suggested improvements:
- train longer (increase epochs)
- tune hyperparameters (learning rate, batch size, etc.)
- use more data
- example: Alpaca with ~52k instruction pairs (vs. 1100 used here)
- try different/larger pretrained backbones than GPT-2 small/medium:
- gpt2 large (~774M)
- gpt2 XL (>1B) if compute allows
- explore parameter-efficient fine-tuning such as LoRA instead of full fine-tuning
9) Final takeaway: the full pipeline and evaluation taxonomy
- Recaps the multi-step instruction pipeline:
- data prep → fine-tune → extract responses → qualitative review → quantitative scoring
- Concludes:
- evaluation is a crucial separate skill
- LLM-as-a-judge is practical for automated scoring
- research remains open due to subjectivity and metric design
Main speakers / sources
- Speaker: The course lecturer (primary narrator; “hello everyone…” and step-by-step walkthrough)
- Sources mentioned:
- MMLU / Massive Multitask Language Understanding paper (general knowledge benchmark)
- Llama 3 Instruct (8B) as the evaluation/judge model
- Ollama for local model inference
- Alpaca dataset for larger instruction tuning (52k)
- LoRA (parameter-efficient fine-tuning concept)