Video summary
The Perfect Local AI GPU? NVIDIA's 5060 Ti 16GB Tested!
Main summary
Key takeaways
Product reviewed
NVIDIA GeForce RTX 5060 Ti (16GB) — positioned as an “MSRP-friendly” local AI GPU for:
- Inference
- Image generation
- Video generation
Key features / setup mentioned
- 16GB VRAM: tests use models that fit fully in GPU memory (no VRAM spilling).
- PCIe Gen 5, X8 width, 128-bit bus:
- Called out as affecting bandwidth-heavy tasks like video.
- Power / design
- 8-pin power connector (preferred over 16-pin).
- Compact / small board (“true two-slot”).
- Expectations: low heat and quiet operation.
- Output ports: 1 HDMI + 3 DisplayPorts
- Software workflows
- Inference: tested via Llama Bench in terminal.
- Image generation: tested using ComfyUI.
- Video generation: also tested (referenced as slow-motion pan; Windows UI/workflow is the default).
Main performance results (unique numerical points)
Inference (Llama Bench; tokens/sec)
Compared against 4090s (reviewer tested “some 4090s” and also references dual 4090 setups).
RTX 5060 Ti (16GB) — prompt vs token generation
-
Qwen / Qwen3-Coder 30B (Q6K + flash attention)
- ~3,570 tokens/sec (prompt processing)
- ~142 tokens/sec (token generation) (Numbers appear in the video subtitles after the 4090 comparison context; the key comparison is later called out explicitly in the verdict section.)
-
Jimma / Jimma 3 12B (Q8)
- ~6,691 prompt tokens/sec
- ~54 tokens/sec token generation
- Mistral Small ~3.2B (Q8)
- ~3,63 prompt tokens/sec
- ~31.7 tokens/sec token generation
Explicit 5060 Ti vs 4090 token generation comparison (best stated numbers)
- ~93 tokens/sec (5060 Ti) vs ~142 tokens/sec (4090)
- For Qwen Coder 30B A3B Q6, sized to fit in 16GB.
Utilization observed (when running 5060 Ti)
- Qwen3-Coder 30B test: utilization around ~47%
- For Jimma 12B: utilization “held well,” but token generation dropped:
- ~24.42 prompt tokens/sec
- ~26 tokens/sec token generation
Image generation (ComfyUI)
- 4090: ~13 seconds per 1024×1024 image
- RTX 5060 Ti: ~42 seconds per similar image
- Described as ~3.2× slower
- Example prompt: “high contrast fluffy black kitten”
- Quality notes:
- Praised attention to window detail, reflections, “aging patina,” and metal framing
- Prompt adherence issue:
- The bird / “crow crackle” ended up inside the window rather than outside (prompt confusion)
Image quality vs prior work
- Reviewer says results are way better than WAN 2.1 (from an earlier video).
- Surprise: results look good despite the ~3× slower generation time.
Video generation (workflow + relative speed)
- RTX 5060 Ti: ~2,220 seconds per video
- Reviewer states this was almost 5× faster on the 4090 side, implying the 4090 ran in roughly ~450 seconds.
- Attribution: PCIe bandwidth differences
- 5060 Ti is X8 (PCIe Gen 5) but with a 128-bit bus, limiting bandwidth for video.
- Usefulness framing:
- For high-level video generation (e.g., WAN 2.2 / 14B), it’s considered usable on faster GPUs.
- WAN 2.2 5B is described as mostly novelty.
Pros (unique points mentioned)
- Strong value at MSRP, and among the “few GPUs available at MSRP.”
- 16GB VRAM “size range” is positioned as worth considering for local AI.
- Newer hardware (reviewer emphasizes demand for new GPUs).
- Good inference performance for price
- Better-than-expected token generation.
- Surprise: a Q6 model size that fully fits VRAM still performs well.
- Compact design likely means low heat and potentially quiet operation.
- 8-pin power connector is praised vs 16-pin.
- Image quality praised (even though slower).
Cons / limitations (unique points mentioned)
- Speed
- Inference: slower than 4090; token generation noticeably lower (e.g., 93 vs 142 tokens/sec).
- Image generation: ~3.2× slower (13s → 42s).
- Video generation: ~5× slower (2,220s per video on 5060 Ti).
- Windows limitation
- Reviewer claims Windows performance is handicapped vs Linux for these workflows.
- Expectation/promise: Linux improvements in a future video.
- Mentions Ulysses and Linux for using more GPUs in video workflows; Windows may require “hoops.”
- Prompt adherence inconsistency
- Bird/crow appeared inside rather than outside in the example.
Comparisons made
RTX 5060 Ti (16GB) vs 4090s
- Inference: 4090 leads; token generation 142 tokens/sec vs 93 tokens/sec (comparable Qwen Coder setup).
- Image generation: 4090 is ~3.2× faster.
- Video generation: 4090 is ~5× faster, attributed to PCIe/bus bandwidth differences.
RTX 5060 Ti vs potential 5090
- Reviewer suggests a 5090 would outperform the 4090 in tested video-generation scenarios (based on bandwidth/architecture expectations).
Windows vs Linux
- Explicit claim: Linux will be ~2.5× faster vs 4090 vs 5060 Ti (reviewer expects Linux gains).
- Linux improvement promised to be shown in upcoming comparisons.
User experience / workflow notes
- ComfyUI used for image generation; defaults to one GPU.
- Example timing behavior:
- First run is slower due to loading; later runs are faster (e.g., cowboy lizard ~21.6s then ~13.23s, trending toward ~14s).
- Video:
- GPU utilization hit 100% and stayed there.
- Render quality described as surprisingly good despite slow generation.
Overall verdict / recommendation (based on the video)
Recommended if you want strong local AI performance at (or near) MSRP—especially for inference—and you’re okay with slower image/video generation.
- Reviewer calls the 5060 Ti 16GB a “great GPU for the price.”
- Not ideal as a primary video-generation GPU compared to 4090/5090, due to large speed gaps driven by bandwidth/power-interface constraints.
- Best results expectation: Linux should improve performance compared to Windows.
Unique points mentioned about the product (consolidated)
- MSRP/value focus; few GPUs available at MSRP
- 16GB VRAM; tests run with models fitting fully in memory
- Compact, two-slot design; expected low heat and quiet
- 8-pin power connector praised
- PCIe Gen 5, X8 width, 128-bit bus
- Inference tested with Llama Bench models (Qwen3-Coder 30B, Jimma 3 12B, Mistral Small)
- Token generation highlighted: 93 tokens/sec (5060 Ti) vs 142 tokens/sec (4090) for comparable Qwen coder setup
- Image speed: 13s (4090) vs 42s (5060 Ti) for ~1024×1024
- Image quality praised; attention to window/detail; compared favorably to WAN 2.1
- Prompt error: bird/crow ended up inside
- Video speed: ~2,220 seconds per clip on 5060 Ti; ~5× slower than 4090
- Video gap explanation: bandwidth limits (PCIe width + bus)
- Windows expected to be worse than Linux; Linux improvements promised
- Notes about single-GPU default in ComfyUI and difficulty using multiple GPUs on Windows
Speakers / viewpoints
- Single main speaker throughout; no clearly differentiated additional speakers.
- Opinions and test results attributed to the same reviewer.