Video summary
Intel just CRUSHED Nvidia & AMD GPU pricing
Main summary
Key takeaways
Product(s) reviewed
- Intel Arc Pro B70 (compared against Arc Pro B50 and multiple competitor GPUs)
- Also benchmarked:
- Intel Arc Pro B60 (referenced as prior testing)
- Nvidia RTX Pro 4000 (Blackwell)
- AMD Radeon AI R9700
- Scaling across four B70 cards
- Image/video tests: only the B70 vs R9700 pair had enough VRAM (32GB) for the chosen image workload.
Key specs / positioning mentioned
Intel Arc Pro B70
- 32 GB VRAM
- Comes in under $1,000
- Requires power (unlike the B50 which draws from PCIe bus only)
- Memory bandwidth cited as “lowest” at 608 GB/s (vs others higher on paper)
Intel Arc Pro B50 (context)
- “Bigger brother” of the B70
- No extra power cables (powered from PCIe bus)
- Previously rated as best bang-for-buck GPU in 2025
Nvidia RTX Pro 4000 (Blackwell)
- 24 GB VRAM (smaller capacity)
- GDDR7
- ~$4,000 retail (as positioned in the video)
- Higher stated bandwidth: 672 GB/s
- Physically “one-slot feel” / very skinny
- Four DisplayPort ports
AMD Radeon AI R9700
- 32 GB VRAM
- ~$1,300, described as ~$350 more per GPU than the B70
- Higher stated bandwidth vs B70: (implied higher than 608)
- Requires/uses ROCm software stack
- Described as very loud (coil whine)
Main features highlighted (what stood out)
- Local AI performance focus
- Benchmarks across LLM inference (prompt processing + token generation)
- Tested with vLLM and llama.cpp stacks
- Multi-GPU scaling
- Tested running four B70s (128GB total VRAM) to assess throughput scaling and bottlenecks
- Workflow support idea
- Emphasis that higher prompt processing helps for larger context and “coding editor” use cases
- Quantization sensitivity
- Tests with different 4-bit quantization methods (notably AWQ vs other approaches)
- Results vary significantly by GPU vendor
- Image + video generation
- ComfyUI for image generation
- LTX2-style video generation
Benchmark / performance findings (numerical highlights)
B60 (prior test, referenced)
llama.cpp / Sickle vs Vulcan
- llama.cpp (Sickle)
- Concurrency 1 (Qwen 34B Q4KM): ~1,010 tok/s prompt processing
- Concurrency 4: prompt processing drops to ~898 tok/s
- Token generation:
- ~66 tok/s (C1)
- ~83 tok/s (C4)
- Vulcan variant
- Prompt processing sometimes higher (e.g., ~1,162 tok/s at C1)
- Token gen worse in some cases (e.g., ~66 vs 44 depending on scenario)
vLLM (best scaling claim for B60)
- Concurrency 1 prompt processing: ~8,118 tok/s
- Concurrency 4 token generation scaling shown as strong:
- ~215 tok/s at C4 token generation
Single-GPU: B70 vs Nvidia RTX Pro 4000 (vLLM)
Concurrency 1 (Qwen 34B BF16 via vLLM)
- Token generation:
- B70 ~56 tok/s
- RTX 4000 ~51 tok/s
- Prompt processing:
- B70 ~12,910 tok/s
- RTX 4000 ~11,745 tok/s
- “Time to first response”
- B70 described as slightly faster (unexpected given RTX’s higher bandwidth)
Concurrency 4
- Token generation:
- B70 ~194 tok/s (Intel top)
- Prompt processing:
- roughly ~12,000 vs ~10–12,000 range mentioned
AWQ quantization behavior
- Under AWQ, results heavily favor Nvidia:
- Concurrency 1 token gen:
- B70 ~72
- RTX 4000 ~89
- Concurrency 4 token gen:
- B70 ~236
- RTX 4000 ~275
- Prompt processing also favors RTX in those AWQ tests (example):
- ~9,825 vs ~11,490
- Concurrency 1 token gen:
Single-GPU: B70 vs AMD Radeon AI R9700 (ROCm stack)
General conclusion
- B70 “destroyed” the R9700 in LLM throughput tests.
Concurrency 1 (Quant 34B, ROCm context)
- Prompt processing:
- B70 ~16,742 (high variance)
- R9700 ~10,800
- Token generation:
- B70 ~56
- R9700 ~43
Concurrency 4
- Prompt processing:
- B70 ~12,564
- R9700 ~9,879
- Token generation:
- B70 ~197
- R9700 ~149
AWQ on AMD (noted failure mode)
- On AMD, AWQ shows particularly poor token generation:
- Prompt processing: R9700 ~13,733 (good)
- Token generation collapses to ~25 tok/s
- B70 under AWQ (same test):
- Token generation ~72 tok/s
- Prompt processing ~9,281
Key interpretation given: software stack maturity is blamed largely for the results (ROCm for AMD; Intel stack also behind Nvidia).
Image generation (large VRAM workload, 32GB-only pair)
- Tested with 1328×1328 images using Quant image model (2512 FP8 quant noted)
- Runtime (time-to-finish):
- R9700: ~133 seconds
- B70: ~147 seconds
- Despite similar outputs, runtime was close (slight advantage to AMD).
- Important caveats:
- Different ComfyUI versions:
- B70: ComfyUI 0.8.2
- R9700: ComfyUI 0.18.1 (newer)
- Intel limited to what’s available in the Intel/Omni package and uses Intel-specific patches/custom nodes; AMD used a ROCm path that “worked a little bit better” for this workload.
- Different ComfyUI versions:
Video generation (LTX2)
- 5-second video, 1280×720
- B70: ~144 seconds
- R9700: ~169 seconds
- Conclusion in that segment: B70 faster.
4× B70 scaling results (128GB total VRAM)
- Model: Qwen 34B AWQ (VL LAM)
Single GPU (AWQ) vs 4 GPUs
Concurrency 1
- Prompt processing roughly doubles:
- ~9,281 → ~18,170 tok/s
- Token generation drops:
- ~72 → ~52 tok/s
Concurrency 4
- Prompt processing increases (example):
- ~10,222 → ~18,000 tok/s
- Token generation decreases:
- ~234 → ~183 tok/s
Bottleneck explanation
- GPU-to-GPU PCIe communication limited to ~63 GB/s even though per-GPU bandwidth is high (608 GB/s stated).
- Limits scaling in token generation on smaller models.
Larger model test (Qwen ~30B instruct / “active 3B parameters”)
- With all four GPUs:
- ~100% compute across devices
- Prompt processing (concurrency 1): ~19,296 tok/s
- Token generation: ~28 tok/s
Coding/editor use case
- Agentic/coding demo attempted; one part failed/stalled.
- Attributed to model age / model not being best and general local-stack limitations.
Pros mentioned about Intel Arc Pro B70
- Very strong value: 32GB under $1,000
- Competitive performance vs AMD in tested LLM workloads (especially outside AWQ-on-Nvidia scenarios)
- Good prompt processing, especially useful for longer context
- Time to first response slightly better than Nvidia in vLLM tests
- Decent multi-card scaling
- Strong prompt throughput gains
- Video generation advantage in the LTX2 test (faster than R9700)
- Benchmarks show it can be nearly as fast as much more expensive Nvidia in some scenarios
Cons / limitations mentioned
- Lower memory bandwidth (608 GB/s) than Nvidia’s cited 672 GB/s, affecting some token generation cases
- Software stack maturity / compatibility is recurring:
- Intel and AMD not on Nvidia’s “level” software-wise
- ROCm described as having a “rocky history”; Intel behind on some model support
- Quantization compatibility varies:
- AWQ results favored Nvidia much more
- AMD token generation could collapse under AWQ
- Consistency / variance
- B70 prompt processing sometimes shows variance vs competitors
- Scaling bottleneck
- PCIe GPU-to-GPU comms limit (~63 GB/s) restricts token generation scaling
- Operational noise / coil whine
- Coil whine on all GPUs; AMD described as very loud
Comparisons made (explicit)
- Intel B70 vs Nvidia RTX Pro 4000
- Similar performance in some vLLM runs, despite RTX’s much higher price and less VRAM
- Nvidia wins more under AWQ quantization tests
- Intel B70 vs AMD Radeon R9700
- B70 generally wins in LLM benchmarks (token generation + concurrency behavior)
- AMD can be closer or win in image generation, but results may be confounded by different ComfyUI versions
- Intel B50 vs B70 context
- B50 praised earlier as best bang-for-buck with simpler PCIe-only power
Unique points list (condensed)
- B50 draws power from PCIe bus; B70 requires power cables
- B70: 32GB VRAM, < $1,000
- B70 single-card comparison supports a 2026 local AI value/performance argument
- Nvidia RTX Pro 4000: 24GB, GDDR7, ~672 GB/s, “one-slot,” 4× DisplayPort, ~2× price
- AMD R9700: 32GB, ~$1,300, ~$350 more than B70 per GPU; ROCm maturity issues
- Main testing uses vLLM for the core professional GPU benchmarks
- vLLM uses memory aggressively for KV cache, filling VRAM per GPU
- Benchmark tool: llama.benchy
- Coil whine observed; differs by model/concurrency
- vLLM Qwen 34B tests:
- B70 slightly better at concurrency 1 vs RTX 4000
- Concurrency 4 improves throughput but generation advantage less consistent
- AWQ works better on Nvidia than on B70 in the tests
- B70 vs R9700 LLM: B70 “destroys” it overall, but variance noted
- Image generation: B70 vs R9700 is close (runtime advantage to AMD), possibly due to ComfyUI version differences
- Intel image workflow uses Intel/Omni limitations and Intel-specific patches/custom nodes
- Video generation: B70 faster (LTX2)
- 4× B70 scaling: prompt throughput scales well; token generation may drop vs single GPU
- Scaling bottleneck: PCIe GPU-to-GPU limited to ~63 GB/s
- Larger model test: prompt throughput strong; GPUs fully utilized
- Coding/editor/agent demo limitations tied to model support/age and stack behavior
- Recommendation implied: verify model/software-stack compatibility; Intel may lag Nvidia for latest models
Verdict / recommendation
- Overall: The Intel Arc Pro B70 is positioned as an exceptional value 32GB GPU for local AI, delivering strong prompt throughput and competitive token generation—especially in the tested vLLM scenarios.
- Caution: Results can swing depending on quantization method (notably AWQ) and software stack support:
- Nvidia may win in some AWQ cases
- Intel/AMD software maturity may limit “latest model” coverage
- Best fit from the video: users who prioritize cost-effective 32GB VRAM, good vLLM performance, and scalable multi-GPU prompt throughput, with willingness to manage software/quantization compatibility.