Video summary
Which AI Models Are Worth Using
Main summary
Key takeaways
Overall premise (what the video is doing)
- The speaker addresses requests for a simple “best vs worst” tier list of AI models.
- They argue a tier list is incomplete because model quality depends on multiple axes, including:
- task type
- real cost
- speed
- performance
- Despite the caveats, they still present a “modern tier list.”
- They use 56 Soul as the baseline reference point, tied to their own experiments and comparisons.
Key reviews / analyses by model (technological + practical takeaways)
56 Soul (baseline reference; top tier discussion)
-
Strengths
- Highly capable for long-running, complex engineering tasks
- Excellent for building applications end-to-end
-
Evidence/examples
- Rewrites a T3 Code mobile app in SwiftUI “entirely” in a single thread
- Compares favorably to 55, with especially notable improvements on long tasks
-
Weakness/quirk
- Logo-finding / multimodal “identification” can be error-prone in their tests
- Often requires user screenshots to correct
-
Positioning
- Recommended as a default model for “most things”
- Not the absolute best for “important code you actually hope to merge,” where another model outranks it
56 Terra (mid/low tier)
-
Pricing
- Cheaper per token than Soul (approximate numbers given)
- But not token-efficient in practice—often consumes more tokens
-
Workflow fit
- Speaker struggles to find a reason to use it
- Measures okay vs Soul per dollar, but doesn’t translate into real workflow preference
-
Tier: C tier
56 Luna (A tier)
-
Core claim
- First “cheap” model in a while that still feels fast, smart, and good for lots of general tasks
-
Best use cases
- Not primarily code-merge work
- Better for “random stuff” in coding tools
-
Agent/tool behavior (in T3 Code)
- Used for:
- title generation / status / category management
- Can:
- make tool calls
- pull and read GitHub/other content
- convert context into useful JSON
- manage threads and generate summaries/context
- Used for:
-
Safety note
- Not trusted for irreversible actions
- Better for reversible context generation/summarization
-
Tier: A tier (speaker heavily uses it)
DeepSeek V4 Flash (value + open weight; around same tier as Luna)
-
Snapshot confusion
- Mentions multiple “V4” snapshots with different capabilities
- Focuses on the newest snapshot
-
Value
- Described as “unbelievable value”
- Considered comparable to Luna (speaker floats putting them in the same tier)
-
Open weight
- Bonus: can run locally / on reasonable hardware
- Enables use cases not possible with closed models
-
Weaknesses
- More prone to loops/distraction
- May need intervention
-
Important limitation
- No vision
- Speaker often relies on image input, so this limitation forces demotion relative to vision-capable options
-
Tier tendency
- “Slightly ahead” due to freedom/open weight
- But the vision gap can push it behind Luna for the speaker’s needs
DeepSeek V4 Pro (F tier due to vision requirement)
-
Complaint
- Does not accept images / lacks vision
- Speaker calls it “embarrassing” given model class/price
-
Tier: F tier
“Kimmy K3” (open weight vision + strong long-task capability; B tier)
-
Strengths
- Surprised the speaker with end-to-end, long complex tasks that touch many components
-
Vision/design/3D
- Strong visual understanding (“pick the right pixels and images”)
- Includes novel 3D capabilities
- Outperforms others on certain 3D tasks
-
Main drawback (cost)
- Not as cheap as many assume
- Expensive enough that (for their settings/workloads) it can be slightly more expensive than 56 Soul
- Example constraint: “max vs high” settings
-
License/pricing constraint analysis
- Open-weight does not guarantee cheap API access
- Pricing is constrained by license terms and revenue caps
- Speaker mentions a ~10M/year revenue threshold and suggests hosting providers may need deals that preserve near-MSRP pricing
-
Tier: B tier
- Impressive, but not top due to cost + tradeoffs
GLM 53 (C tier-ish; “catch-up” feel + no vision)
-
Progress vs 52
- Better at staying on task
- Fewer stupid loops
-
Framing
- Feels like refinement/catch-up, not a brand-new leap
-
Limitation
- No vision
- Leads to harsher ranking
-
Tier: C tier
Composer 2.5 / Composer 2.x lineage (D tier)
-
Background
- Cursor / “SpaceX’s cursor” model line built from Kimmy variants
- Heavy RL fine-tuning to be a fast coding model
-
Reality check
- Speaker doesn’t use it much; availability limitations matter
-
Limitations called out
- Removed from Grock Build and other tools (not broadly available)
- Only available in certain products (not over API)
- Performance/price tiers are misleading
- Fast option too expensive relative to claims
-
Tier: D tier (maybe “high D” in demo contexts)
Grock 46 (lower than Grock 45; utility but not preferred)
-
Regression
- Worse token efficiency
- Slower and less aligned with “fast as hell” expectations
-
Orchestration tradeoff
- Grock 45 tracked multiple disparate tasks better via sub-agent orchestration
- Grock 46 costs more tokens to do similar orchestration
-
Competition / comparison
- Speaker prefers 56 Soul on low
- Similar quality at slightly faster response for comparable price
-
Where it fits
- Might be useful in Grockbot for assistance
- Goal: keep reasoning effort low
-
Tier
- Around B tier originally considered
- But implies it’s lower than 45
- Still useful, just not preferred
Muse Spark (next to Grock 46; underrated processing agent)
-
Strength
- Fast and surprisingly accurate for non-code workflow processing
- Example: triaging open PRs and deciding priorities
-
Performance comparison
- Finishes in under 2 minutes
- Other models: 10+ minutes or even hours
-
Code quality
- Speaker dislikes the code it writes
- Not preferred for coding changes
-
Pricing/support
- Very cheap “contributor” tier mentioned
- Expects open-weight soon
-
Tier
- Implied comparable to Grock 46 (near it), not top
Gemini models (mostly F tier; “Google tier” at bottom)
Gemini 3.7 Flash → F tier
-
Speaker sentiment
- Strongly dislikes it; described as “really bad”
-
Blame
- Says Google’s direction is bad
- Performance/cost inefficiencies vs older Flash models
Main cost/efficiency critique (core technical analysis)
- Earlier Flash (e.g., Gemini 20 Flash) was extremely cheap and fast
- Later versions introduced “thinking”/reasoning behavior that massively increases:
- output token cost
- and/or token usage in real workflows
- Real-world cost increases claimed to be 10–100x vs older Flash, worsening over time
Token-efficiency reasoning example
- Speaker estimates Gemini Flash would consume very large token counts
- Example: hundreds of thousands at higher reasoning tiers
-
Conclusion: poor value
-
Pro series also criticized
- Faster tokens/sec doesn’t help
- Because the model generates far more tokens overall
-
Net recommendation: “Just don’t use Gemini.”
Sonnet 5 (low D / mostly F tier depending on environment)
-
Use case
- Testing Claude Code sub-agents inside T3 Code
- Visualizable math-ish tasks
- Can spin up sub agents
-
But
- Not token efficient
- More expensive than it should be
-
Hard rule
- Don’t pick it for anything except Cloud Code sub-agents
-
Tier: “Low D,” with strict usage constraints
Opus 5 (C / high D due to “mimic behavior”)
-
Initial impression
- Seems capable of matching Fable-level detail
- Catches things other models miss
-
Failure mode during merging
- When code must actually run/merge:
- outputs become vague/jargony
- code execution quality is poor
- When code must actually run/merge:
-
Behavior critique
- “Looks like a duck… mimic behavior”
- It convinces you it’s smart until you try to integrate it
-
Tier
- Low C capability but high D cost/behavior
- Net ends up around high D
Anthropic Fable 5 (winner; S tier—top of chart)
-
Tier
- S tier
- Framed as “the only S tier model right now”
-
Strengths
- Best overall code quality and depth of understanding
- Most willing to merge output
- Strong at:
- double-checking work
- handling deep/hard engineering exploration
-
How it differs from Soul (workflow + agentic behavior)
- Soul is “thorough but writes worse code”
- Fable understands goals better and produces code the speaker feels safer merging
- With Fable/Soul, the speaker can give vague instructions and approve
- Models then handle:
- computer-use verification
- sub-agent review before PR creation
-
Tradeoffs / imperfections
- Trips over itself sometimes
- Loses track occasionally
- Takes unnecessary shortcuts
- Touches things it shouldn’t
- Some issues attributed to Anthropic-side constraints
-
Pricing/business note
- Speaker says they would never pay full API price
- Relies on sub discounts (they “burn” a lot of subscription value weekly)
Notable “guides/tips” embedded in the tier logic
- Use Luna/Flash as background context processors rather than irreversible “merge code” agents.
- Treat non-vision models as unacceptable if you frequently paste images (explicit demotion rule applied to V4 Flash and V4 Pro).
- Choose reasoning tiers based on token efficiency:
- higher reasoning can explode token counts and reduce cost-value
- Prefer models that fit your agent workflow:
- If you rely on iterative decomposition + sub-agent orchestration, some models work better
- If you want “tell me what to build and it will PR,” the best models differ fundamentally
Sponsors / product feature mentioned (technology context)
- Browserbase
- Described as an agent-friendly browser solution that provides web access for tasks like:
- searching like Google
- fetching known URLs
- full interactive browsing (including sign-in/click actions)
- Used to argue that models without web context feel less capable
- Browser access improves agent intelligence
- Described as an agent-friendly browser solution that provides web access for tasks like:
Main speakers/sources
- Main speaker: The solo narrator/reviewer (no specific name provided in the subtitles)
- Primary sources/models discussed:
- 56 Soul, 56 Terra, 56 Luna
- DeepSeek V4 Flash/Pro
- Kimmy K3
- GLM 53
- Composer 2.5 (Composer 2.x lineage)
- Grock 46
- Muse Spark
- Gemini (various Flash/Pro variants)
- Anthropic Sonnet 5
- Anthropic Opus 5
- Anthropic Fable 5
- Sponsor source: Browserbase
- Also mentioned: a hiring sponsor, G2I