Video summary
Fastest Qwen 3.8 27B in Llama.cpp? DFlash 2 + n-gram Explained & Benchmarked!
Main summary
Key takeaways
Main topic
This video benchmarks Qwen 3.8 (27B) running in llama.cpp with speculative decoding, focusing on D-Flash 2 (the drafter technique) and how it can be combined with n-gram (engram) speculative decoding via llama.cpp’s speculative/verification pipeline.
The author’s central claim is that D-Flash 2 can substantially increase the number of accepted tokens while adding only a small latency cost, and that mixing speculative methods can improve throughput on limited hardware.
Key product/algorithm features explained (D-Flash 2)
Goal
- Improve acceptance rate (more drafted tokens survive verification)
- Keep added overhead low
Core mechanism: “path selection”
D-Flash 2 uses a path selection strategy:
- Keep the top 16 candidates per position inside the drafter.
- Score neighboring pairs, considering how a candidate (e.g., B) follows prior candidate(s) (e.g., A), to select a more coherent multi-token trajectory.
- This is intended to address a limitation of D-Flash v1, where positions were drafted in parallel and couldn’t coordinate consistency between adjacent tokens.
Claimed cost/benefit
- Roughly +20% accepted tokens per verification pass
- About ~1% extra cycle latency
- The author notes some claims depend on how llama.cpp is implemented/configured and the exact algorithm details.
Quantified results / benchmark takeaways
Accepted-length improvement (author’s measurements)
- Example improvement: acceptance length rises from approximately 4.27 → 4.71 using D-Flash 2 path selection.
- The video states this achieves acceptance gains without requiring a major increase in model size compared to a heavier alternative (referred to as “park” in subtitles—likely another speculative method/model variant).
“Oracle” comparisons / evaluation methodology
The evaluation includes metrics such as:
- Recall/oracle-like acceptance: whether the drafted token matches the correct token (under an oracle framing)
- A stronger oracle-style variant: if the correct token exists anywhere in the candidate set, acceptance can be higher
Reported observations:
- Drafting correctness improves with the 16-candidate set
- e.g., ~99% “exists in candidates” versus lower top-1 correctness.
Mean acceptance vs latency tradeoff (dataset-dependent)
Across multiple datasets (e.g., GSM8K, MAF500, and HumanEval/MBP/MT-Bench-like sets):
- Mean acceptance improves with D-Flash 2
- The video claims approximately:
- ~+21% mean acceptance length over baseline D-Flash
- ~small latency overhead (≈ +1%)
- Some tasks reach high acceptance for later tokens in the draft block.
D-Flash 2 internal architecture (why it’s faster)
Drafter-only changes
- D-Flash 2 primarily modifies the drafter-only portion
- Verification/model-side behavior is described as essentially unchanged
Convolution “backbone” integration
The drafter adds lightweight convolution to help drafting condition better on predecessor information, even though tokens are drafted in parallel:
- Uses a two-tap dynamic convolution idea (as described in the explanation)
- Intended benefits:
- Better conditioning on predecessor hidden states
- Reduced accuracy drop toward the end of draft blocks (improves the “tail” of drafted length)
Attention/MLP unchanged, convolution inserted around sublayers
- Convolution doesn’t replace attention
- It is inserted before/after attention and MLP within drafter sublayers
- The convolution is dynamic:
- mixing weights depend on current hidden state (content-dependent), not a fixed rule
Combining speculative methods: D-Flash 2 + engram (n-gram)
Key insight from the tutorial
- In llama.cpp, you don’t have to choose only one speculative method.
- The author tests:
- D-Flash 2 alone
- Engram alone
- Combinations, where llama.cpp can route/choose the applicable speculative method
Findings
As reported:
- On a real multi-turn coding-style benchmark (Live CodeBench variant + synthetic harness), adding the engram/lookup component can yield extra throughput
- “% tokens per second increase” varies
- Effectiveness is workload-dependent
- Some synthetic story-generation cases behave oddly
- Some results show small differences or even negative changes
Benchmark setup + tuning guidance
Hardware/test environment
- Server: “6,000 Blackwell” GPU / 96 GB RAM
- llama.cpp concurrency=1 for main comparisons
- Uses greedy sampling for reproducibility
Quantization
- Notes different quantization formats for target vs drafter models
- Exact matching may not always be possible across versions
Benchmark suite categories
-
Live CodeBench
- competitive programming style
- generally low-context / multitopic
-
Custom multi-coding benchmark (Gradio app)
- designed to mimic multi-turn repo coding / agentic workflows
-
Synthetic harness
- controls throughput/acceptance across generation lengths
- caution: may overfit/inflate results due to repeated/fake story behavior
Draft length tuning
- D-Flash 2 draft length (in this version): up to 7 draft tokens simultaneously
- The author experiments with draft-per-block and observes an example peak around 5 in one sweep
Parameter mention: draft_perm / early stopping drafter (capacity control)
The author discusses a parameter (referred to as draft_perm in subtitles) to stop drafting early when the drafter is unsure.
- Expected benefit:
- more relevant under high concurrency (more requests → more GPU pressure)
- reducing wasted drafting compute
- Observations:
- in their setup, the effect was hard to see
- MTP reportedly showed clearer impact than D-Flash
Limitations / cautions
- No deep context decode measurements in this study
- Multi-turn benchmarks may contain artifacts if prompts repeat or reuse earlier outputs
- No deep correlation study between “quality” and acceptance beyond the author’s statement that acceptance isn’t degraded (based on earlier tests/smoke checks)
- Benchmarks should be rerun for confidence (“next day” verification); the author mentions cooling/noise checks and repeats where feasible
Main speakers/sources
- Speaker: Bukash (channel host; described as a senior inference engineer)
- Referenced ecosystem/components:
- llama.cpp
- Qwen 3.8 27B
- D-Flash 2
- n-gram/engram speculative decoding
- speculative variants such as MTP (as named in subtitles)