Video summary
How Fast Can One RTX 3060 Actually Run 35B (llama.cpp enhancement)?
Main summary
Key takeaways
Summary of Technological Concepts, Features, and Findings
- Target task & hardware claim: The video argues that an RTX 3060 can run a ~35B MoE model (via an llama.cpp enhancement) at ~70 tokens/sec using a Q4 model, producing token-for-token identical output to the baseline (same quantization, same results).
- Core performance goal: Reduce the overhead of offloading MoE “experts” between GPU VRAM and system RAM over PCIe. When experts aren’t resident on the GPU, the system must swap/load weights, causing major latency.
What Was Changed (Key Design Ideas)
1. Replace “layer offload” with “slot-based caching”
- Traditional approach: When the model doesn’t fit in VRAM, it offloads experts by layer, treating each layer as a block.
- Proposed approach: Allocate a fixed number of “slots” on the GPU; when the router selects an expert:
- if the expert is in a slot, reuse it,
- if it isn’t, execute it elsewhere,
- evict/replace using a cache policy (conceptually similar to LRU).
2. The “zero-missing” trick (the enabler)
Modify expert execution so that:
- If the requested expert is not on the GPU, it returns zeros instead of running the expert there.
Then:
- Run one pass over experts on the GPU,
- Run another pass for experts in RAM,
- Sum the results to match the original output exactly (missing experts contribute zero on the GPU pass, so RAM provides the real values).
3. Work scheduling/plumbing matters (big real bottleneck)
- An early/incorrect performance version arranged GPU/CPU work in an alternating pattern.
- The scheduler regrouped work in a way that caused expert weights to bounce across PCIe repeatedly (even though cache logic was “correct”).
- Fix: ensure GPU work runs in one contiguous run, CPU work in another, then combine once—eliminating unnecessary synchronization/switching overhead.
- Outcome: performance jumped from ~6 tok/s back to ~44 tok/s, close to the ~42 tok/s baseline.
Cache Findings (Why the Stats Looked Like It Shouldn’t Help)
The surprising observation
- Expert usage statistics appeared flat (no small “hot set”), suggesting caching shouldn’t help.
Why it worked anyway
- The distribution is flat over long time, but at any given moment only a handful of experts are active.
- A cache targeting recently used experts (without requiring model-specific hotness knowledge) achieved ~45% hit rate.
- An “oracle” fixed best set based on perfect knowledge of the whole run achieved only ~41%.
Conclusion
- The right caching question is: “what is hot right now”, not “what is popular overall.”
MoE-Specific Performance Analysis Numbers
-
PCIe transfer cost dominates when swapping is frequent
- Some cache designs were slower due to constant swapping.
- One bad variant implied ~127 MB transferred per token (effectively disastrous).
- A winning variant swapped far less (~9 MB), showing that on PCIe it’s loading/moving, not just “miss counting,” that kills performance.
-
Best working principle
- Choose experts once before running, so you “pay loading cost” up front.
- Avoid moving expert weights during token generation.
Additional Optimizations Layered on Top
4. Slot count sweep (find how much fits)
- Increasing slots improved speed until VRAM ran out.
- The model used 256 experts per layer (not 128 as assumed).
- On the RTX 3060, about half of experts fit, enabling most gains.
5. Interaction with speculative decoding (MTP/speculative)
-
Combining:
- baseline speculative decoding improvements, and
- expert caching underneath it, produced even larger gains.
-
Claimed results:
- speculative alone: ~55 tok/s (from ~45 tok/s),
- with cache: up to ~70 tok/s.
Reasoning: speculative decoding checks many guessed tokens, increasing expert demand; having more relevant experts already cached makes verification cheap and the extra guesses worthwhile.
6. Concurrency: run GPU and CPU expert chains in parallel
- Originally GPU and CPU parts ran sequentially.
- New design runs them concurrently (separate CPU thread while GPU continues).
- Claimed improvement: 70–75 tok/s, because GPU spends less time waiting for CPU.
7. Batch-size guardrail
- Overlap/concurrency helps only for small prompt batches.
- For large batches, overlap/merge cost outweighs savings, so it turns off automatically to preserve correctness and performance.
Practical Guide: How to Run It (Profile-Based Setup)
Offline profiling step
- Use a separate tool to capture which experts the router selects during generation.
- Run with two prompt styles:
- “codeish”
- “chatty”
- Generate a few hundred tokens per profile (recommended ~500 tokens between replications; described as enough).
- Merge profiles into a single file.
Serving step
- Run the normal
llama.cppcommand plus two flags:- point to the profile,
- set “how many slots per layer” to use.
- On startup, check logs for:
- how many layers/slots were set up,
- how much data was uploaded.
- If the expected log line is missing, the system likely fell back to baseline—often because the requested slots didn’t fit in VRAM.
Limitations / Expectations Depending on VRAM
- Gains depend on the fraction of experts that fit:
- RTX 3060 + 35B model: ~half experts fit → strong results.
- Larger “Laguna” (118B) example: only 36/256 experts fit → gain drops to ~5%.
- Rule of thumb from the video:
- below about 15–20% fit, fixed caching stops being the right approach.
- motivates future work: repicking experts periodically (e.g., every 500 tokens) rather than only once.
Main Speakers / Sources
- Primary speaker: The video’s single narrator/author (implied by first-person references such as “I built it,” “I traced it,” “I ran the sweep,” “my fork”).
- Referenced components/systems:
- llama.cpp, including its scheduler,
- MoE router/expert selection logic,
- mentions speculative decoding / MTP.