Video summary
Local AI Video and Image Generation on the Intel's Low Cost 32GB GPU!
Main summary
Key takeaways
Overview
Linus Tech Tips demonstrates running local AI video and image generation using an Intel-powered setup built around a 32GB Intel B760 “GPU” (32GB VRAM) connected to a Minisforum PC via an OcuLink dock. The main focus is how well recent generation models run locally (on-desk hardware) using ComfyUI as the workflow interface.
What’s demonstrated: Local video generation
Core workflow (ComfyUI + Intel-optimized setup)
- Uses ComfyUI, a node/workflow “wrapper” for models, with an Intel-optimized configuration.
Model used
- Runs Minimax H3 for text-to-video generation.
Model placement / performance behavior
- Most of the model fits into the GPU’s VRAM.
- Some parts offload to system memory.
Typical generation speed
- Short 5-second clips are a “sweet spot.”
- One timed scenario (prompt/clip change) took ~2 min 36 sec total.
- Longer clips are possible, but take more time.
Resolution / quality tradeoff
- Example output shown around 480p (~0.4 megapixels).
- Mentions potential to go up to roughly 1080p, but performance degrades as output size increases.
Prompting + template usage
- Uses ComfyUI templates, including a Super Mario Bros universe prompt example.
- Also generates videos from scratch with custom prompts (e.g., a spoof ad and character scenes).
Example outputs and observed limitations
Example outputs
- McDonald’s-style parody commercial (with audio).
- Captain Picard scene(s), iterated through multiple prompts/scripts.
Observed limitation: dialogue handling in longer clips
- In an approximately 10-second clip, voices between two characters can become mixed up, even when the characters visibly interact.
Motion graphics / looping backgrounds
- Can generate looping backgrounds when instructed to align first/last frames.
- If prompts are vague, it may produce gibberish/incorrect text.
- A “cheat” workaround is mentioned: using a blank placeholder at loop endpoints.
Still-image features (image-to-video + bridging)
The model is used to:
- Convert a single still image into video (e.g., an animated character scene with music).
- Use an ending image to create a transition/merge (“bridge”) between “before” and “after” frames.
Node-based workflow (more than simple prompting)
- Highlights ComfyUI’s node-based interface, enabling more complex pipelines than a single prompt box.
- Supports different operation orderings and wired components.
Image generation tests on the same system
ComfyUI image generation acceleration
- ComfyUI is also used to run image generation accelerated on the Intel device.
Retro McDonald’s + timing
- Demonstrates a retro McDonald’s image, then a Star Wars X-Wing variant.
- Timed run: 1024×1024 image generated in about 35 seconds.
Flux 2 Klein 9B (from Hugging Face)
- Another image model tested: Flux 2 (Flux 2 Klein 9B).
- Notes:
- Uses open weights, but is not open source.
- Requires Hugging Face contact registration and an API key.
- Mentions likely watermarking/tracing back to the user.
Storyboard example (constraint following)
- Flux 2 Klein 9B is used to create a three-panel storyboard:
- A female engineer repairing a robot.
- Same woman and same clothes across panels.
- Upscales to 1280×720 and claims direction-following was reasonably good.
Review / analysis takeaways: why local matters
- Main benefit: iteration without cloud token limits
- Cloud services (e.g., Google VEO) can cap experimentation via monthly tokens.
- Local runs allow many prompt iterations, often reaching usable results in ~2–5 minutes depending on clip length.
- Suggested practical workflow:
- Iterate locally with a faster/cheaper setup.
- Optionally send the refined prompt to a cloud model for higher quality.
- Hardware bottleneck:
- VRAM is emphasized as the key limiting factor.
- His comment suggests even an 8GB GPU may still run decent local LLMs for work.
- Purchase advice:
- Don’t rush—start with what you have, and expect ongoing efficiency improvements.
Installation / tooling notes
- Installation is described as a “moving target” due to frequent model updates.
- Uses Frontier Models to help configure local models via Codex CLI.
- Runs on Linux (Ubuntu 26.04) for best compatibility.
- Uses ChatGPT as a troubleshooting/configuration helper, especially for Intel GPU quirks.
- Notes ComfyUI Intel optimizations may lag behind newly released models:
- Example: a newer LTX 2.5 release that fits the card’s memory, but ComfyUI wasn’t updated yet to fully use it.
Main speakers / sources
- Linus Tech Tips (Linus / Linus Seidman) — primary speaker and experimenter.
- ComfyUI — node/workflow framework for generation pipelines.
- Minimax H3 — local video generation model demonstrated.
- Flux 2 Klein 9B — local image generation model demonstrated.
- Hugging Face — hosting platform for open-weight models and API key requirement for the Flux model.