Video summary

Fastest Qwen 3.8 27B in Llama.cpp? DFlash 2 + n-gram Explained & Benchmarked!

Main summary

Key takeaways

Technology

Main topic

This video benchmarks Qwen 3.8 (27B) running in llama.cpp with speculative decoding, focusing on D-Flash 2 (the drafter technique) and how it can be combined with n-gram (engram) speculative decoding via llama.cpp’s speculative/verification pipeline.

The author’s central claim is that D-Flash 2 can substantially increase the number of accepted tokens while adding only a small latency cost, and that mixing speculative methods can improve throughput on limited hardware.


Key product/algorithm features explained (D-Flash 2)

Goal

  • Improve acceptance rate (more drafted tokens survive verification)
  • Keep added overhead low

Core mechanism: “path selection”

D-Flash 2 uses a path selection strategy:

  • Keep the top 16 candidates per position inside the drafter.
  • Score neighboring pairs, considering how a candidate (e.g., B) follows prior candidate(s) (e.g., A), to select a more coherent multi-token trajectory.
  • This is intended to address a limitation of D-Flash v1, where positions were drafted in parallel and couldn’t coordinate consistency between adjacent tokens.

Claimed cost/benefit

  • Roughly +20% accepted tokens per verification pass
  • About ~1% extra cycle latency
  • The author notes some claims depend on how llama.cpp is implemented/configured and the exact algorithm details.

Quantified results / benchmark takeaways

Accepted-length improvement (author’s measurements)

  • Example improvement: acceptance length rises from approximately 4.27 → 4.71 using D-Flash 2 path selection.
  • The video states this achieves acceptance gains without requiring a major increase in model size compared to a heavier alternative (referred to as “park” in subtitles—likely another speculative method/model variant).

“Oracle” comparisons / evaluation methodology

The evaluation includes metrics such as:

  • Recall/oracle-like acceptance: whether the drafted token matches the correct token (under an oracle framing)
  • A stronger oracle-style variant: if the correct token exists anywhere in the candidate set, acceptance can be higher

Reported observations:

  • Drafting correctness improves with the 16-candidate set
    • e.g., ~99% “exists in candidates” versus lower top-1 correctness.

Mean acceptance vs latency tradeoff (dataset-dependent)

Across multiple datasets (e.g., GSM8K, MAF500, and HumanEval/MBP/MT-Bench-like sets):

  • Mean acceptance improves with D-Flash 2
  • The video claims approximately:
    • ~+21% mean acceptance length over baseline D-Flash
    • ~small latency overhead (≈ +1%)
  • Some tasks reach high acceptance for later tokens in the draft block.

D-Flash 2 internal architecture (why it’s faster)

Drafter-only changes

  • D-Flash 2 primarily modifies the drafter-only portion
  • Verification/model-side behavior is described as essentially unchanged

Convolution “backbone” integration

The drafter adds lightweight convolution to help drafting condition better on predecessor information, even though tokens are drafted in parallel:

  • Uses a two-tap dynamic convolution idea (as described in the explanation)
  • Intended benefits:
    • Better conditioning on predecessor hidden states
    • Reduced accuracy drop toward the end of draft blocks (improves the “tail” of drafted length)

Attention/MLP unchanged, convolution inserted around sublayers

  • Convolution doesn’t replace attention
  • It is inserted before/after attention and MLP within drafter sublayers
  • The convolution is dynamic:
    • mixing weights depend on current hidden state (content-dependent), not a fixed rule

Combining speculative methods: D-Flash 2 + engram (n-gram)

Key insight from the tutorial

  • In llama.cpp, you don’t have to choose only one speculative method.
  • The author tests:
    • D-Flash 2 alone
    • Engram alone
    • Combinations, where llama.cpp can route/choose the applicable speculative method

Findings

As reported:

  • On a real multi-turn coding-style benchmark (Live CodeBench variant + synthetic harness), adding the engram/lookup component can yield extra throughput
    • “% tokens per second increase” varies
  • Effectiveness is workload-dependent
    • Some synthetic story-generation cases behave oddly
    • Some results show small differences or even negative changes

Benchmark setup + tuning guidance

Hardware/test environment

  • Server: “6,000 Blackwell” GPU / 96 GB RAM
  • llama.cpp concurrency=1 for main comparisons
  • Uses greedy sampling for reproducibility

Quantization

  • Notes different quantization formats for target vs drafter models
  • Exact matching may not always be possible across versions

Benchmark suite categories

  1. Live CodeBench

    • competitive programming style
    • generally low-context / multitopic
  2. Custom multi-coding benchmark (Gradio app)

    • designed to mimic multi-turn repo coding / agentic workflows
  3. Synthetic harness

    • controls throughput/acceptance across generation lengths
    • caution: may overfit/inflate results due to repeated/fake story behavior

Draft length tuning

  • D-Flash 2 draft length (in this version): up to 7 draft tokens simultaneously
  • The author experiments with draft-per-block and observes an example peak around 5 in one sweep

Parameter mention: draft_perm / early stopping drafter (capacity control)

The author discusses a parameter (referred to as draft_perm in subtitles) to stop drafting early when the drafter is unsure.

  • Expected benefit:
    • more relevant under high concurrency (more requests → more GPU pressure)
    • reducing wasted drafting compute
  • Observations:
    • in their setup, the effect was hard to see
    • MTP reportedly showed clearer impact than D-Flash

Limitations / cautions

  • No deep context decode measurements in this study
  • Multi-turn benchmarks may contain artifacts if prompts repeat or reuse earlier outputs
  • No deep correlation study between “quality” and acceptance beyond the author’s statement that acceptance isn’t degraded (based on earlier tests/smoke checks)
  • Benchmarks should be rerun for confidence (“next day” verification); the author mentions cooling/noise checks and repeats where feasible

Main speakers/sources

  • Speaker: Bukash (channel host; described as a senior inference engineer)
  • Referenced ecosystem/components:
    • llama.cpp
    • Qwen 3.8 27B
    • D-Flash 2
    • n-gram/engram speculative decoding
    • speculative variants such as MTP (as named in subtitles)

Original video