Video summary
Intel is saving the GPU market...kind of
Main summary
Key takeaways
Product reviewed
Intel Arc Pro B70 (“Battlemage”) workstation GPU — 32GB VRAM The video compares it mainly to other workstation/AI GPUs, with a strong focus on AI performance and homelab/self-hosting features.
Key features mentioned
- 32GB GDDR6 VRAM (positioned as the standout value)
- Two-slot, blower-style design
- Single 8-pin power connector
- Key specs called out
- ~600 GB/s memory bandwidth
- Base clock: 2.2 GHz
- TDP: 230 W
- Price: just under $1,000 (for 32GB)
- Native SR-IOV support (major differentiator for homelab/self-hosters)
- Splits the physical GPU into virtual GPUs (vGPUs) assigned to separate VMs/containers
- Claimed to be not paywalled/enterprise-locked by Intel (unlike some competitors)
- Example from the creator: configured into four 8GB vGPUs, visible in
lspcias separate devices
- AI backend note
- Mentions a “proper back end” in SYCL (driver/software ecosystem still immature overall vs NVIDIA)
Performance / user experience (from the video)
AI (main focus)
- LlamaBench with SYCL using models in the 32GB VRAM tier:
- Qwen 3 4B
- Llama 3.1 8B
- Mistral 7B
- Qwen 3 32B
- Single-stream results
- Prefill (prompt processing) reported as:
- Over 1,000 tokens/sec prefill
- Over 1,500 tokens/sec on Qwen 3 4B
- Decode/streaming usability guidance
- The creator notes “usable” real-time chat/coding assistants typically need ~20–30 tokens/sec decode
- Qwen 3 32B falls just under ~20 tokens/sec decode → feels slower but “usable” (and potentially more accurate/helpful)
- Prefill (prompt processing) reported as:
- Concurrency scaling
- Concurrency test with Llama.cpp
- Throughput (overall tokens/sec) scales up until ~8 streams, then plateaus (possibly dips)
- The creator suggests better concurrency may improve when “VL/stack becomes more stable on Intel cards”
- Latency (TTFT)
- TTFT with 1 stream: 0.068 seconds
- At 16 streams: ~0.25 seconds
- The creator argues regular users won’t notice much unless workloads are heavier
Gaming / general GPU use
- 3DMark: score described as “pretty reasonable” (no exact number given)
- Cyberpunk test
- 1080p, high settings
- XeSS Super Resolution 2.0 (auto) + ray tracing
- Reported ~100 FPS
- Creator caveat
- XeSS/ray tracing is “how most people play,” but acknowledged it could be seen as “cheating”
- Gaming recommendation conclusion
- Not recommended for pure gamers, largely due to Intel being new and older games having issues
Pros (explicitly emphasized)
- Excellent value for VRAM at the workstation/AI level:
- 32GB for just under $1,000
- ~$31.25 per GB
- Native SR-IOV enables GPU partitioning for home labs without enterprise/paywall restrictions
- AI performance described as “pretty good/solid”, especially for 4B–8B class models in this VRAM tier
- Works in virtualized setups:
- vGPU passed to Windows VM for Time Spy
- vGPU passed to Ubuntu VM for AI workloads
- Pricing/value comparison favors the B70 (see below)
Cons / limitations (explicitly mentioned)
- Software/driver maturity
- Gaming and AI support is said to be less mature than competitors
- For pure gaming, older titles may have compatibility/performance issues
- Big model decoding speed
- Qwen 3 32B decode is just under the creator’s “comfort” range (may feel slower for real-time chat/coding)
- Concurrency scaling is limited and depends on stack stability (plateau around ~8 streams)
Comparisons made (pricing and capacity)
Value comparison ($ per GB VRAM)
- Intel Arc Pro B70 (32GB): ~$31.25 per GB
- Compared against:
- RTX 5090 / “Radeon AI Pro 9700” pricing mentioned
- Creator’s stated figures:
- RTX 5090: ~$125 per GB
- Radeon AI Pro 9700: ~$43.75 per GB
- Conclusion: B70 comes out ahead on cost efficiency (most VRAM per dollar in this tier, per the video)
Within the Battlemage lineup
- B60 (24GB)
- Less Xe processors
- Slower memory bandwidth
- 24GB VRAM
- B50
- Creator already has one
- ~70W power draw
- Mentioned as lower-end / cheaper option (exact price not provided)
Unique points / claims collected (all distinct product-related points)
- B70 is positioned as a top option for certain users, especially AI/home-lab/self-hosters.
- 32GB VRAM at under $1,000 is highlighted as unusually good value (VRAM per dollar).
- Two-slot blower design with single 8-pin power.
- Native SR-IOV is a differentiator for virtualization and vGPU splitting.
- SR-IOV is described as freely available (not locked behind enterprise tiers/paywalls).
- Demonstrated vGPU partitioning into 4×8GB and mentions optional 2×16GB configurations.
- vGPUs appear as separate devices (via
lspci), simplifying VM assignment. - AI performance tested using LlamaBench (SYCL) with multiple models in the 32GB tier.
- Prefill speeds: >1,000 tokens/sec overall; >1,500 on Qwen 3 4B.
- Decode usability: Qwen 3 32B falls just under ~20 tokens/sec decode (creator’s threshold range).
- Concurrency: throughput rises up to ~8 streams then plateaus.
- TTFT latency is very low (0.068s at 1 stream; ~0.25s at 16).
- Gaming test shows ~100 FPS in Cyberpunk at 1080p high with XeSS SR 2.0 + ray tracing.
- Recommendation: not for pure gamers, but recommended for homelabbers and AI workstation users on budget.
- Broader note: Intel’s relative immaturity may cause issues, especially in older games.
Speakers / perspectives
- Only one main speaker is present (the creator “Brett” per subtitles).
- No other distinct speakers are providing separate evaluations beyond the creator’s own commentary.
Overall verdict (concise)
Recommended primarily for homelabbers and AI/self-hosters, mainly because the 32GB VRAM value and especially native SR-IOV for easy vGPU partitioning are the core strengths. Not the best pick for pure gamers, due to Intel’s newer GPU ecosystem and potential older-game/software immaturity.