Video summary

Evaluating fine-tuned LLM using Ollama

Main summary

Key takeaways

Technology

Summary (tech concepts, product features, evaluation/guide details)

This lecture continues an instruction fine-tuning project for a small LLM and then focuses on evaluating instruction-tuned models.

1) Fine-tuning recap (stages 1–2)

Stage 1: Data preparation

  • Uses an instruction dataset of ~1100 instruction input/output pairs.
  • Builds training utilities:
    • downloading
    • batching
    • creating data loaders
  • Uses token-ID batching logic so batches have consistent token lengths.
  • Dataset split:
    • 85% train
    • 10% test
    • 5% validation

Stage 2: Fine-tuning

  • Loads a pretrained GPT-2 checkpoint: 355M parameters (often referred to as “gpt2 medium”).
  • Fine-tunes on the instruction dataset so the model learns improved instruction-following.
  • Shows training/validation loss behavior.
  • Due to CPU compute limits, runs for 1 epoch
    • Suggestion: 2+ epochs (with better hardware) typically improves results.

2) Why evaluation is needed (stage 3: evaluation)

Even after fine-tuning, the model can produce semantically wrong but plausible responses.

Example failure (active → passive):

  • Model: “prepared by the chef…”
  • Ground truth: “cooked by the chef…”

The lecture emphasizes that evaluation isn’t just a simple yes/no label task. Instead, you often need to:

  • compare generated text vs. ground truth
  • apply a scoring system for response quality

3) Step 6: Extract and save model responses for the test set

After fine-tuning, the code:

  • runs the model on the entire test subset
  • decodes token IDs back to text
  • removes the repeated prompt/instruction portion
  • collects outputs for later analysis

It also demonstrates qualitative checks on 3 test samples, showing:

  • sometimes outputs are close/mostly correct
  • sometimes outputs are off-topic or incorrect
  • many errors are attributed to training for only 1 epoch

4) Create an evaluation-ready dataset file

The lecture builds a new JSON file: instruction_data_with_response.json, including:

  • instruction
  • input
  • true output
  • model response

This format supports later automated scoring across many items.

5) Save the fine-tuned model

To avoid re-training, checkpoint saving is highlighted:

  • save with torch.save(model_state_dict)
  • later restore with load_state_dict

Example saved filename resembles: gpt2-medium-355M...pth.

6) Step 7: Quantitative evaluation using another LLM + Ollama/AMA

The lecture outlines three evaluation approaches for instruction-tuned LLMs:

  1. Benchmarking general knowledge using MLU (Massive Multitask Language Understanding)
    • Example: MMLU uses 57 tasks across domains (STEM, humanities, etc.).
  2. Human preference comparison
    • humans rate competing model outputs
  3. LLM-as-a-judge
    • a strong instruction model scores responses vs. ground truth
    • the lecture implements this approach (#3)

How scoring is automated (LLM-as-judge)

  • Uses Llama 3 Instruct (8B) as the judge model.
  • Runs locally via:
    • Ollama (model hosting/inference on a laptop)
    • “AMA” (mentioned as an inference application/interface; text generation only)

Install/run examples:

  • macOS/Linux: ollama run llama3
  • Windows alternative: ollama serve then ollama run llama3

Python helper concept:

  • query_model(prompt, model) to send prompts to Ollama-served Llama 3 and retrieve scores.

Scoring prompt behavior

  • The judge is prompted with:
    • instruction/input
    • ground-truth output
    • model response
  • The judge returns a score, typically requested on a 0–100 scale.
  • The lecture demonstrates:
    • first requesting rationale
    • then modifying the prompt to request integer-only scores

Compute constraints

  • CPU evaluation is slow and memory-heavy:
    • prompting many examples can cause laptop hangs
    • full test-set evaluation may be impractical
  • Suggestion: use GPU or higher RAM.

7) Results and interpretation

For 3 demo items scored by the judge:

  • One response gets a high partial score (~85) despite style differences
  • One response gets low (~20) because it doesn’t answer correctly (e.g., thunderstorm cloud type)
  • One response gets zero because the output is factually wrong (e.g., Pride and Prejudice author)

Notes:

  • Scores may not be perfectly deterministic (minor variation possible).

8) Reported average performance + improvement directions

  • Average score is ~just above 50
    • context: limited training data, 1 epoch, CPU constraints
  • Observes common issues:
    • incorrect copying of output text
    • failure to follow the intended instruction

Suggested improvements:

  • train longer (increase epochs)
  • tune hyperparameters (learning rate, batch size, etc.)
  • use more data
    • example: Alpaca with ~52k instruction pairs (vs. 1100 used here)
  • try different/larger pretrained backbones than GPT-2 small/medium:
    • gpt2 large (~774M)
    • gpt2 XL (>1B) if compute allows
  • explore parameter-efficient fine-tuning such as LoRA instead of full fine-tuning

9) Final takeaway: the full pipeline and evaluation taxonomy

  • Recaps the multi-step instruction pipeline:
    • data prep → fine-tune → extract responses → qualitative review → quantitative scoring
  • Concludes:
    • evaluation is a crucial separate skill
    • LLM-as-a-judge is practical for automated scoring
    • research remains open due to subjectivity and metric design

Main speakers / sources

  • Speaker: The course lecturer (primary narrator; “hello everyone…” and step-by-step walkthrough)
  • Sources mentioned:
    • MMLU / Massive Multitask Language Understanding paper (general knowledge benchmark)
    • Llama 3 Instruct (8B) as the evaluation/judge model
    • Ollama for local model inference
    • Alpaca dataset for larger instruction tuning (52k)
    • LoRA (parameter-efficient fine-tuning concept)

Original video