Video summary

Intel just CRUSHED Nvidia & AMD GPU pricing

Main summary

Key takeaways

Product Review

Product(s) reviewed

  • Intel Arc Pro B70 (compared against Arc Pro B50 and multiple competitor GPUs)
  • Also benchmarked:
    • Intel Arc Pro B60 (referenced as prior testing)
    • Nvidia RTX Pro 4000 (Blackwell)
    • AMD Radeon AI R9700
    • Scaling across four B70 cards
  • Image/video tests: only the B70 vs R9700 pair had enough VRAM (32GB) for the chosen image workload.

Key specs / positioning mentioned

Intel Arc Pro B70

  • 32 GB VRAM
  • Comes in under $1,000
  • Requires power (unlike the B50 which draws from PCIe bus only)
  • Memory bandwidth cited as “lowest” at 608 GB/s (vs others higher on paper)

Intel Arc Pro B50 (context)

  • “Bigger brother” of the B70
  • No extra power cables (powered from PCIe bus)
  • Previously rated as best bang-for-buck GPU in 2025

Nvidia RTX Pro 4000 (Blackwell)

  • 24 GB VRAM (smaller capacity)
  • GDDR7
  • ~$4,000 retail (as positioned in the video)
  • Higher stated bandwidth: 672 GB/s
  • Physically “one-slot feel” / very skinny
  • Four DisplayPort ports

AMD Radeon AI R9700

  • 32 GB VRAM
  • ~$1,300, described as ~$350 more per GPU than the B70
  • Higher stated bandwidth vs B70: (implied higher than 608)
  • Requires/uses ROCm software stack
  • Described as very loud (coil whine)

Main features highlighted (what stood out)

  • Local AI performance focus
    • Benchmarks across LLM inference (prompt processing + token generation)
    • Tested with vLLM and llama.cpp stacks
  • Multi-GPU scaling
    • Tested running four B70s (128GB total VRAM) to assess throughput scaling and bottlenecks
  • Workflow support idea
    • Emphasis that higher prompt processing helps for larger context and “coding editor” use cases
  • Quantization sensitivity
    • Tests with different 4-bit quantization methods (notably AWQ vs other approaches)
    • Results vary significantly by GPU vendor
  • Image + video generation
    • ComfyUI for image generation
    • LTX2-style video generation

Benchmark / performance findings (numerical highlights)

B60 (prior test, referenced)

llama.cpp / Sickle vs Vulcan

  • llama.cpp (Sickle)
    • Concurrency 1 (Qwen 34B Q4KM): ~1,010 tok/s prompt processing
    • Concurrency 4: prompt processing drops to ~898 tok/s
    • Token generation:
      • ~66 tok/s (C1)
      • ~83 tok/s (C4)
  • Vulcan variant
    • Prompt processing sometimes higher (e.g., ~1,162 tok/s at C1)
    • Token gen worse in some cases (e.g., ~66 vs 44 depending on scenario)

vLLM (best scaling claim for B60)

  • Concurrency 1 prompt processing: ~8,118 tok/s
  • Concurrency 4 token generation scaling shown as strong:
    • ~215 tok/s at C4 token generation

Single-GPU: B70 vs Nvidia RTX Pro 4000 (vLLM)

Concurrency 1 (Qwen 34B BF16 via vLLM)

  • Token generation:
    • B70 ~56 tok/s
    • RTX 4000 ~51 tok/s
  • Prompt processing:
    • B70 ~12,910 tok/s
    • RTX 4000 ~11,745 tok/s
  • “Time to first response”
    • B70 described as slightly faster (unexpected given RTX’s higher bandwidth)

Concurrency 4

  • Token generation:
    • B70 ~194 tok/s (Intel top)
  • Prompt processing:
    • roughly ~12,000 vs ~10–12,000 range mentioned

AWQ quantization behavior

  • Under AWQ, results heavily favor Nvidia:
    • Concurrency 1 token gen:
      • B70 ~72
      • RTX 4000 ~89
    • Concurrency 4 token gen:
      • B70 ~236
      • RTX 4000 ~275
    • Prompt processing also favors RTX in those AWQ tests (example):
      • ~9,825 vs ~11,490

Single-GPU: B70 vs AMD Radeon AI R9700 (ROCm stack)

General conclusion

  • B70 “destroyed” the R9700 in LLM throughput tests.

Concurrency 1 (Quant 34B, ROCm context)

  • Prompt processing:
    • B70 ~16,742 (high variance)
    • R9700 ~10,800
  • Token generation:
    • B70 ~56
    • R9700 ~43

Concurrency 4

  • Prompt processing:
    • B70 ~12,564
    • R9700 ~9,879
  • Token generation:
    • B70 ~197
    • R9700 ~149

AWQ on AMD (noted failure mode)

  • On AMD, AWQ shows particularly poor token generation:
    • Prompt processing: R9700 ~13,733 (good)
    • Token generation collapses to ~25 tok/s
  • B70 under AWQ (same test):
    • Token generation ~72 tok/s
    • Prompt processing ~9,281

Key interpretation given: software stack maturity is blamed largely for the results (ROCm for AMD; Intel stack also behind Nvidia).


Image generation (large VRAM workload, 32GB-only pair)

  • Tested with 1328×1328 images using Quant image model (2512 FP8 quant noted)
  • Runtime (time-to-finish):
    • R9700: ~133 seconds
    • B70: ~147 seconds
  • Despite similar outputs, runtime was close (slight advantage to AMD).
  • Important caveats:
    • Different ComfyUI versions:
      • B70: ComfyUI 0.8.2
      • R9700: ComfyUI 0.18.1 (newer)
    • Intel limited to what’s available in the Intel/Omni package and uses Intel-specific patches/custom nodes; AMD used a ROCm path that “worked a little bit better” for this workload.

Video generation (LTX2)

  • 5-second video, 1280×720
    • B70: ~144 seconds
    • R9700: ~169 seconds
  • Conclusion in that segment: B70 faster.

4× B70 scaling results (128GB total VRAM)

  • Model: Qwen 34B AWQ (VL LAM)

Single GPU (AWQ) vs 4 GPUs

Concurrency 1

  • Prompt processing roughly doubles:
    • ~9,281 → ~18,170 tok/s
  • Token generation drops:
    • ~72 → ~52 tok/s

Concurrency 4

  • Prompt processing increases (example):
    • ~10,222 → ~18,000 tok/s
  • Token generation decreases:
    • ~234 → ~183 tok/s

Bottleneck explanation

  • GPU-to-GPU PCIe communication limited to ~63 GB/s even though per-GPU bandwidth is high (608 GB/s stated).
  • Limits scaling in token generation on smaller models.

Larger model test (Qwen ~30B instruct / “active 3B parameters”)

  • With all four GPUs:
    • ~100% compute across devices
    • Prompt processing (concurrency 1): ~19,296 tok/s
    • Token generation: ~28 tok/s

Coding/editor use case

  • Agentic/coding demo attempted; one part failed/stalled.
  • Attributed to model age / model not being best and general local-stack limitations.

Pros mentioned about Intel Arc Pro B70

  • Very strong value: 32GB under $1,000
  • Competitive performance vs AMD in tested LLM workloads (especially outside AWQ-on-Nvidia scenarios)
  • Good prompt processing, especially useful for longer context
    • Time to first response slightly better than Nvidia in vLLM tests
  • Decent multi-card scaling
    • Strong prompt throughput gains
  • Video generation advantage in the LTX2 test (faster than R9700)
  • Benchmarks show it can be nearly as fast as much more expensive Nvidia in some scenarios

Cons / limitations mentioned

  • Lower memory bandwidth (608 GB/s) than Nvidia’s cited 672 GB/s, affecting some token generation cases
  • Software stack maturity / compatibility is recurring:
    • Intel and AMD not on Nvidia’s “level” software-wise
    • ROCm described as having a “rocky history”; Intel behind on some model support
    • Quantization compatibility varies:
      • AWQ results favored Nvidia much more
      • AMD token generation could collapse under AWQ
  • Consistency / variance
    • B70 prompt processing sometimes shows variance vs competitors
  • Scaling bottleneck
    • PCIe GPU-to-GPU comms limit (~63 GB/s) restricts token generation scaling
  • Operational noise / coil whine
    • Coil whine on all GPUs; AMD described as very loud

Comparisons made (explicit)

  • Intel B70 vs Nvidia RTX Pro 4000
    • Similar performance in some vLLM runs, despite RTX’s much higher price and less VRAM
    • Nvidia wins more under AWQ quantization tests
  • Intel B70 vs AMD Radeon R9700
    • B70 generally wins in LLM benchmarks (token generation + concurrency behavior)
    • AMD can be closer or win in image generation, but results may be confounded by different ComfyUI versions
  • Intel B50 vs B70 context
    • B50 praised earlier as best bang-for-buck with simpler PCIe-only power

Unique points list (condensed)

  1. B50 draws power from PCIe bus; B70 requires power cables
  2. B70: 32GB VRAM, < $1,000
  3. B70 single-card comparison supports a 2026 local AI value/performance argument
  4. Nvidia RTX Pro 4000: 24GB, GDDR7, ~672 GB/s, “one-slot,” 4× DisplayPort, ~2× price
  5. AMD R9700: 32GB, ~$1,300, ~$350 more than B70 per GPU; ROCm maturity issues
  6. Main testing uses vLLM for the core professional GPU benchmarks
  7. vLLM uses memory aggressively for KV cache, filling VRAM per GPU
  8. Benchmark tool: llama.benchy
  9. Coil whine observed; differs by model/concurrency
  10. vLLM Qwen 34B tests:
    • B70 slightly better at concurrency 1 vs RTX 4000
    • Concurrency 4 improves throughput but generation advantage less consistent
  11. AWQ works better on Nvidia than on B70 in the tests
  12. B70 vs R9700 LLM: B70 “destroys” it overall, but variance noted
  13. Image generation: B70 vs R9700 is close (runtime advantage to AMD), possibly due to ComfyUI version differences
  14. Intel image workflow uses Intel/Omni limitations and Intel-specific patches/custom nodes
  15. Video generation: B70 faster (LTX2)
  16. 4× B70 scaling: prompt throughput scales well; token generation may drop vs single GPU
  17. Scaling bottleneck: PCIe GPU-to-GPU limited to ~63 GB/s
  18. Larger model test: prompt throughput strong; GPUs fully utilized
  19. Coding/editor/agent demo limitations tied to model support/age and stack behavior
  20. Recommendation implied: verify model/software-stack compatibility; Intel may lag Nvidia for latest models

Verdict / recommendation

  • Overall: The Intel Arc Pro B70 is positioned as an exceptional value 32GB GPU for local AI, delivering strong prompt throughput and competitive token generation—especially in the tested vLLM scenarios.
  • Caution: Results can swing depending on quantization method (notably AWQ) and software stack support:
    • Nvidia may win in some AWQ cases
    • Intel/AMD software maturity may limit “latest model” coverage
  • Best fit from the video: users who prioritize cost-effective 32GB VRAM, good vLLM performance, and scalable multi-GPU prompt throughput, with willingness to manage software/quantization compatibility.

Original video