Video summary
Qwen 3.8 27b x NInfer = 1.5x Faster@256K Context(on one 24GB 4090)
Main summary
Key takeaways
Video summary (tech concepts, product features, analysis)
The video compares Qwen 3.8 27B inference speed and long-context performance using a new inference engine/framework called NInfer, versus llama.cpp (with mention of MTP). It also contrasts NInfer’s design with vLLM and SGLang, focusing on differences in engineering philosophy and caching/concurrency strategies.
What NInfer is (core idea)
NInfer is described as an inference engine written “from zero”, engineered to be extremely optimized for a single fixed execution path, rather than broad compatibility.
- It locks to one model family (primarily Qwen 3.8) and one hardware target
- In the described setup, it’s effectively a 4090-specific fork (“14 90 fork”), optimized for:
- Qwen 3.8 27B + RTX 4090
How it differs from other engines
- llama.cpp: prioritizes broad compatibility (e.g., many model formats such as GGUF and many backends). The tradeoff is potential compromises/bugs that can appear under specific conditions.
- vLLM: focuses on high throughput under concurrency using techniques like:
- PagedAttention
- continuous batching
- prefix caching
- SGLang: focuses on cache reuse across workflows using:
- RadixAttention, which builds prefixes into a tree to avoid recomputation—particularly useful for agent-like repetitive behavior.
Key framing: NInfer is positioned as “pure hardware juicing” (like a Formula 1 car on one track), not a general-purpose engine.
Deployment experience described
The creator describes moving from llama.cpp to NInfer through a tuning-and-testing cycle:
- Previously deployed Qwen 3.8 27B on llama.cpp
- Spent dozens of hours tuning to reach approximately:
- ~70–80 tokens/sec at ~200k context
- Spent dozens of hours tuning to reach approximately:
- After adopting NInfer:
- Used DeepSeek V4 Flash to run about:
- 5 hours of configuration/metrics testing for NInfer parameters
- Produced a best configuration, then compared results against llama.cpp
- Used DeepSeek V4 Flash to run about:
Reported performance results (RTX 4090, 24GB)
Context size / memory fit
- NInfer can run up to 256K (≈262K) native context on a 24GB RTX 4090.
- At 256K, the video claims about ~1.25GB of VRAM remains, enabling additional stacking (e.g., running a vision model alongside).
KV cache approach
The video describes Qwen 3.8 as using a hybrid architecture, enabling large context on 24GB because:
- KV cache is quantized, not full precision
- Mentioned approach: RK4V4-E8
- described as “rotating jumping compression”
- intended to save VRAM with only small precision loss
It further claims NInfer’s optimization reduces KV cache demands enough for 256K to fit in 24GB.
Tokens/sec numbers
- Reported average generation speed:
- ~100 tokens/sec
- In “real use,” the creator reports:
- ~130 tokens/sec for small tasks
- just over ~100 tokens/sec for long chains
Relative speed claim vs llama.cpp + MTP
- Estimated ~1.5× speedup
- llama.cpp often lands around ~50–70 tok/s, depending on conditions
Cold-start / prefill trade-off
The creator notes a downside with NInfer: cold prefill is worse.
- First token latency (TTFT / cold request):
- llama.cpp: ~5 seconds
- NInfer: ~12–13 seconds for cold starts
- Prefill throughput trade-off (cold-start case):
- not “2× slower”
- roughly ~20–30% slower prefill throughput
After warm-up and caching behavior improves, performance is said to improve, with TTFT dropping to a few seconds later.
Power / cost analysis (measured)
Power draw
- Highest-load wall power (during decoding):
- ~530–540W (NInfer setup)
- During preview/cache hits:
- ~200–300W
Power savings and electricity estimate
- Estimated power savings:
- ~25–30%
- Electricity usage estimate:
- Old (llama.cpp): ~6–7 kWh/day
- New (NInfer): ~4.5–5 kWh/day
- Annualized electricity cost estimate:
- ~$140/year
The creator also claims qualitative benefits tied to measurements:
- less heat
- fans spin less
Stability / operational behavior
Startup validation
NInfer is described as more stable in production because it:
- validates capacity at startup
- if you exceed budget, it doesn’t start rather than crashing mid-run
- avoids mid-session OOM behavior
Comparison to llama.cpp
Llama.cpp is described as potentially less stable on long sessions due to KV slot management differences, including:
- mismatches between unified KV behavior and expectations
- default slot settings (e.g., “default is four”)
- needing to pin KV slots to avoid instability as context grows
Limitations mentioned
- Main downside (beyond cold-start latency):
- Concurrency trade-off
- The creator wants full 256K context and is concerned that achieving more concurrency would require giving up context.
- The video frames this as acceptable because:
- concurrency on a single GPU with extremely long context isn’t the goal
- priority is speed and long-chain correctness, not multi-request throughput
Overall conclusion from the creator
- The creator moved their production setup to NInfer.
- Claims include:
- speed improvements made generation more healthy and usable
- no quality drop in their tests
- fewer trouble-free sessions than llama.cpp
They also mention broader benchmarking:
- DeepSeek V4.x “Flash” is still described as even faster
- but NInfer is preferred for the creator’s specific needs: long-context performance, stability, and their 4090 setup
Main speakers / sources
- Primary speaker: the YouTube video author/host (creator/operator describing their own deployment and measurements)
- Referenced tools/engines:
- NInfer
- llama.cpp
- vLLM
- SGLang
- DeepSeek V4 Flash (used to test/config-search)
- Model referenced:
- Qwen 3.8 27B (and Qwen 3.6 27B mentioned historically)
Rate this summary
Your feedback will help improve summaries.
Improve this summary
Reprocess with a stronger model when the summary feels incomplete or inaccurate.
Translate summary in another language
Ask questions to this video
Chat for follow-up questions, clarifications, and source-backed answers.