Video summary
Local AI Is Only 4 Months Behind. I Did the Math.
Main summary
Key takeaways
Summary
Ash compares downloadable open-weight AI models with closed commercial models using the Artificial Analysis Intelligence Index, a composite score out of 100 covering coding, multi-step agent tasks, general knowledge, and science. Because older models are scored on the same index version, he argues that their scores can be compared directly.
Three hardware tiers
-
One used RTX 3090 — about $1,250
- The video runs Alibaba’s Qwen 3.8 27B model in a roughly 17 GB, four-bit file. It fits in the card’s 24 GB of memory and is reported to generate about 40 tokens per second.
- The model scores 34, ahead of Google’s Gemini 3.1 Pro (30) and level with GPT-5.6 Terra on its high setting.
- Ash sees this as the best value for everyday writing, questions, and coding assistance. It is less capable at long, independent jobs across large codebases.
- He notes that a newer RTX 5090 is faster, while Nvidia’s GB10-based DGX Spark is slower on this model. He attributes this to fast memory mattering more than large capacity for a model of this size. A Mac with 32 GB can also hold the model, though performance varies.
-
Gaming PC with a 12 GB GPU and 64 GB RAM — about $2,500
- A free engine called Strata runs Alibaba’s much larger Qwen 3.8 Flash-Next 125B model by placing the most active parts on the GPU and using regular system memory for the rest.
- The engine’s author reports 50–90 tokens per second, depending on the hardware and model version.
- The model scores 40, above GPT-6 Luna (38), which Ash describes as the smallest GPT-6 model and the one offered to free ChatGPT users.
- Ash’s verdict: it is worthwhile as an upgrade if you already own a gaming PC, but building one solely for AI may take years to recoup compared with a subscription.
-
Two Mac Studios — about $19,000
- Xiaomi’s MiMo 2.6 Pro is presented as the best downloadable model, scoring 46—the same as GPT-6 Astra on its low setting and Grok 4.7.
- MiMo is around one trillion parameters and free under the MIT license, but its large memory requirements make local use difficult. The video says there was no published real-world test of it on a Mac at the time.
- As a currently runnable alternative, GLM 5.3 scores 45. Its roughly 343 GB, three-bit version can be split across two 256 GB Mac Studios. Ash could not find a clean speed test of that exact setup and declines to estimate its speed.
- Ash says local hosting can be valuable for organizations that cannot send code or data outside their premises. He also cites an estimated task cost of about $0.13 for hosted MiMo versus $0.82 for GPT-6 Astra on low.
Broader findings and caveats
- The video places Claude Opus 5.5 at the top of the scoreboard with 58, while GPT-6 Astra and Gemini 4 score 53. Ash says the gap between the best open model and the frontier is concentrated in difficult, long, multi-step work.
- Epoch AI research is cited as estimating that open models trail closed models by about four months on average, up from three months in its earlier study. Nathan Lambert of the Allen Institute is cited as estimating a three-to-five-month gap.
- The leading open models discussed are from Chinese labs. Ash contrasts them with lower-scoring American open models, while noting that models run locally do not send data to an external service.
- For the hardest tasks, Ash recommends renting access to frontier models rather than relying on local hardware. He cites Claude Opus 5.5 on maximum setting as taking about 11 minutes before its first output and averaging about $6 per task on the referenced tests.
- Scores, prices, and availability reflect the video’s comparisons and may change as models and benchmarks are updated.
Reviews, guides, and tutorials referenced
- Earlier RAM video: Discussed the memory shortage and its effects on hardware prices.
- Strata explainer video: Covered how Strata divides a large model between GPU memory and system memory, as well as the economics of building a PC for it.
- Video on Mistral Large 4: Ash says he had covered its preview, which scored 38; the transcript also mentions a planned open-weight release.
These are references to other videos, not step-by-step tutorials included in this one.
Main speaker and sources
- Main speaker: Ash, presenting his own comparison and conclusions.
- Sources cited: Artificial Analysis Intelligence Index and speed tests; Epoch AI research; Nathan Lambert of the Allen Institute; and Strata’s author for reported generation speeds.
Rate this summary
Your feedback will help improve summaries.
Improve this summary
Reprocess with a stronger model when the summary feels incomplete or inaccurate.
Translate summary in another language
Ask questions to this video
Chat for follow-up questions, clarifications, and source-backed answers.