Video summary

How Fast Can One RTX 3060 Actually Run 35B (llama.cpp enhancement)?

Main summary

Key takeaways

Technology

Summary of Technological Concepts, Features, and Findings

  • Target task & hardware claim: The video argues that an RTX 3060 can run a ~35B MoE model (via an llama.cpp enhancement) at ~70 tokens/sec using a Q4 model, producing token-for-token identical output to the baseline (same quantization, same results).
  • Core performance goal: Reduce the overhead of offloading MoE “experts” between GPU VRAM and system RAM over PCIe. When experts aren’t resident on the GPU, the system must swap/load weights, causing major latency.

What Was Changed (Key Design Ideas)

1. Replace “layer offload” with “slot-based caching”

  • Traditional approach: When the model doesn’t fit in VRAM, it offloads experts by layer, treating each layer as a block.
  • Proposed approach: Allocate a fixed number of “slots” on the GPU; when the router selects an expert:
    • if the expert is in a slot, reuse it,
    • if it isn’t, execute it elsewhere,
    • evict/replace using a cache policy (conceptually similar to LRU).

2. The “zero-missing” trick (the enabler)

Modify expert execution so that:

  • If the requested expert is not on the GPU, it returns zeros instead of running the expert there.

Then:

  • Run one pass over experts on the GPU,
  • Run another pass for experts in RAM,
  • Sum the results to match the original output exactly (missing experts contribute zero on the GPU pass, so RAM provides the real values).

3. Work scheduling/plumbing matters (big real bottleneck)

  • An early/incorrect performance version arranged GPU/CPU work in an alternating pattern.
  • The scheduler regrouped work in a way that caused expert weights to bounce across PCIe repeatedly (even though cache logic was “correct”).
  • Fix: ensure GPU work runs in one contiguous run, CPU work in another, then combine once—eliminating unnecessary synchronization/switching overhead.
  • Outcome: performance jumped from ~6 tok/s back to ~44 tok/s, close to the ~42 tok/s baseline.

Cache Findings (Why the Stats Looked Like It Shouldn’t Help)

The surprising observation

  • Expert usage statistics appeared flat (no small “hot set”), suggesting caching shouldn’t help.

Why it worked anyway

  • The distribution is flat over long time, but at any given moment only a handful of experts are active.
  • A cache targeting recently used experts (without requiring model-specific hotness knowledge) achieved ~45% hit rate.
  • An “oracle” fixed best set based on perfect knowledge of the whole run achieved only ~41%.

Conclusion

  • The right caching question is: “what is hot right now”, not “what is popular overall.”

MoE-Specific Performance Analysis Numbers

  • PCIe transfer cost dominates when swapping is frequent

    • Some cache designs were slower due to constant swapping.
    • One bad variant implied ~127 MB transferred per token (effectively disastrous).
    • A winning variant swapped far less (~9 MB), showing that on PCIe it’s loading/moving, not just “miss counting,” that kills performance.
  • Best working principle

    • Choose experts once before running, so you “pay loading cost” up front.
    • Avoid moving expert weights during token generation.

Additional Optimizations Layered on Top

4. Slot count sweep (find how much fits)

  • Increasing slots improved speed until VRAM ran out.
  • The model used 256 experts per layer (not 128 as assumed).
  • On the RTX 3060, about half of experts fit, enabling most gains.

5. Interaction with speculative decoding (MTP/speculative)

  • Combining:

    • baseline speculative decoding improvements, and
    • expert caching underneath it, produced even larger gains.
  • Claimed results:

    • speculative alone: ~55 tok/s (from ~45 tok/s),
    • with cache: up to ~70 tok/s.

Reasoning: speculative decoding checks many guessed tokens, increasing expert demand; having more relevant experts already cached makes verification cheap and the extra guesses worthwhile.

6. Concurrency: run GPU and CPU expert chains in parallel

  • Originally GPU and CPU parts ran sequentially.
  • New design runs them concurrently (separate CPU thread while GPU continues).
  • Claimed improvement: 70–75 tok/s, because GPU spends less time waiting for CPU.

7. Batch-size guardrail

  • Overlap/concurrency helps only for small prompt batches.
  • For large batches, overlap/merge cost outweighs savings, so it turns off automatically to preserve correctness and performance.

Practical Guide: How to Run It (Profile-Based Setup)

Offline profiling step

  • Use a separate tool to capture which experts the router selects during generation.
  • Run with two prompt styles:
    • “codeish”
    • “chatty”
  • Generate a few hundred tokens per profile (recommended ~500 tokens between replications; described as enough).
  • Merge profiles into a single file.

Serving step

  • Run the normal llama.cpp command plus two flags:
    • point to the profile,
    • set “how many slots per layer” to use.
  • On startup, check logs for:
    • how many layers/slots were set up,
    • how much data was uploaded.
  • If the expected log line is missing, the system likely fell back to baseline—often because the requested slots didn’t fit in VRAM.

Limitations / Expectations Depending on VRAM

  • Gains depend on the fraction of experts that fit:
    • RTX 3060 + 35B model: ~half experts fit → strong results.
    • Larger “Laguna” (118B) example: only 36/256 experts fit → gain drops to ~5%.
  • Rule of thumb from the video:
    • below about 15–20% fit, fixed caching stops being the right approach.
    • motivates future work: repicking experts periodically (e.g., every 500 tokens) rather than only once.

Main Speakers / Sources

  • Primary speaker: The video’s single narrator/author (implied by first-person references such as “I built it,” “I traced it,” “I ran the sweep,” “my fork”).
  • Referenced components/systems:
    • llama.cpp, including its scheduler,
    • MoE router/expert selection logic,
    • mentions speculative decoding / MTP.

Original video