Video summary

Which AI Models Are Worth Using

Main summary

Key takeaways

Technology

Overall premise (what the video is doing)

  • The speaker addresses requests for a simple “best vs worst” tier list of AI models.
  • They argue a tier list is incomplete because model quality depends on multiple axes, including:
    • task type
    • real cost
    • speed
    • performance
  • Despite the caveats, they still present a “modern tier list.”
  • They use 56 Soul as the baseline reference point, tied to their own experiments and comparisons.

Key reviews / analyses by model (technological + practical takeaways)

56 Soul (baseline reference; top tier discussion)

  • Strengths

    • Highly capable for long-running, complex engineering tasks
    • Excellent for building applications end-to-end
  • Evidence/examples

    • Rewrites a T3 Code mobile app in SwiftUI “entirely” in a single thread
    • Compares favorably to 55, with especially notable improvements on long tasks
  • Weakness/quirk

    • Logo-finding / multimodal “identification” can be error-prone in their tests
    • Often requires user screenshots to correct
  • Positioning

    • Recommended as a default model for “most things”
    • Not the absolute best for “important code you actually hope to merge,” where another model outranks it

56 Terra (mid/low tier)

  • Pricing

    • Cheaper per token than Soul (approximate numbers given)
    • But not token-efficient in practice—often consumes more tokens
  • Workflow fit

    • Speaker struggles to find a reason to use it
    • Measures okay vs Soul per dollar, but doesn’t translate into real workflow preference
  • Tier: C tier


56 Luna (A tier)

  • Core claim

    • First “cheap” model in a while that still feels fast, smart, and good for lots of general tasks
  • Best use cases

    • Not primarily code-merge work
    • Better for “random stuff” in coding tools
  • Agent/tool behavior (in T3 Code)

    • Used for:
      • title generation / status / category management
    • Can:
      • make tool calls
      • pull and read GitHub/other content
      • convert context into useful JSON
      • manage threads and generate summaries/context
  • Safety note

    • Not trusted for irreversible actions
    • Better for reversible context generation/summarization
  • Tier: A tier (speaker heavily uses it)


DeepSeek V4 Flash (value + open weight; around same tier as Luna)

  • Snapshot confusion

    • Mentions multiple “V4” snapshots with different capabilities
    • Focuses on the newest snapshot
  • Value

    • Described as “unbelievable value”
    • Considered comparable to Luna (speaker floats putting them in the same tier)
  • Open weight

    • Bonus: can run locally / on reasonable hardware
    • Enables use cases not possible with closed models
  • Weaknesses

    • More prone to loops/distraction
    • May need intervention
  • Important limitation

    • No vision
    • Speaker often relies on image input, so this limitation forces demotion relative to vision-capable options
  • Tier tendency

    • “Slightly ahead” due to freedom/open weight
    • But the vision gap can push it behind Luna for the speaker’s needs

DeepSeek V4 Pro (F tier due to vision requirement)

  • Complaint

    • Does not accept images / lacks vision
    • Speaker calls it “embarrassing” given model class/price
  • Tier: F tier


“Kimmy K3” (open weight vision + strong long-task capability; B tier)

  • Strengths

    • Surprised the speaker with end-to-end, long complex tasks that touch many components
  • Vision/design/3D

    • Strong visual understanding (“pick the right pixels and images”)
    • Includes novel 3D capabilities
    • Outperforms others on certain 3D tasks
  • Main drawback (cost)

    • Not as cheap as many assume
    • Expensive enough that (for their settings/workloads) it can be slightly more expensive than 56 Soul
    • Example constraint: “max vs high” settings
  • License/pricing constraint analysis

    • Open-weight does not guarantee cheap API access
    • Pricing is constrained by license terms and revenue caps
    • Speaker mentions a ~10M/year revenue threshold and suggests hosting providers may need deals that preserve near-MSRP pricing
  • Tier: B tier

    • Impressive, but not top due to cost + tradeoffs

GLM 53 (C tier-ish; “catch-up” feel + no vision)

  • Progress vs 52

    • Better at staying on task
    • Fewer stupid loops
  • Framing

    • Feels like refinement/catch-up, not a brand-new leap
  • Limitation

    • No vision
    • Leads to harsher ranking
  • Tier: C tier


Composer 2.5 / Composer 2.x lineage (D tier)

  • Background

    • Cursor / “SpaceX’s cursor” model line built from Kimmy variants
    • Heavy RL fine-tuning to be a fast coding model
  • Reality check

    • Speaker doesn’t use it much; availability limitations matter
  • Limitations called out

    • Removed from Grock Build and other tools (not broadly available)
    • Only available in certain products (not over API)
    • Performance/price tiers are misleading
    • Fast option too expensive relative to claims
  • Tier: D tier (maybe “high D” in demo contexts)


Grock 46 (lower than Grock 45; utility but not preferred)

  • Regression

    • Worse token efficiency
    • Slower and less aligned with “fast as hell” expectations
  • Orchestration tradeoff

    • Grock 45 tracked multiple disparate tasks better via sub-agent orchestration
    • Grock 46 costs more tokens to do similar orchestration
  • Competition / comparison

    • Speaker prefers 56 Soul on low
    • Similar quality at slightly faster response for comparable price
  • Where it fits

    • Might be useful in Grockbot for assistance
    • Goal: keep reasoning effort low
  • Tier

    • Around B tier originally considered
    • But implies it’s lower than 45
    • Still useful, just not preferred

Muse Spark (next to Grock 46; underrated processing agent)

  • Strength

    • Fast and surprisingly accurate for non-code workflow processing
    • Example: triaging open PRs and deciding priorities
  • Performance comparison

    • Finishes in under 2 minutes
    • Other models: 10+ minutes or even hours
  • Code quality

    • Speaker dislikes the code it writes
    • Not preferred for coding changes
  • Pricing/support

    • Very cheap “contributor” tier mentioned
    • Expects open-weight soon
  • Tier

    • Implied comparable to Grock 46 (near it), not top

Gemini models (mostly F tier; “Google tier” at bottom)

Gemini 3.7 Flash → F tier

  • Speaker sentiment

    • Strongly dislikes it; described as “really bad”
  • Blame

    • Says Google’s direction is bad
    • Performance/cost inefficiencies vs older Flash models

Main cost/efficiency critique (core technical analysis)

  • Earlier Flash (e.g., Gemini 20 Flash) was extremely cheap and fast
  • Later versions introduced “thinking”/reasoning behavior that massively increases:
    • output token cost
    • and/or token usage in real workflows
  • Real-world cost increases claimed to be 10–100x vs older Flash, worsening over time

Token-efficiency reasoning example

  • Speaker estimates Gemini Flash would consume very large token counts
  • Example: hundreds of thousands at higher reasoning tiers
  • Conclusion: poor value

  • Pro series also criticized

    • Faster tokens/sec doesn’t help
    • Because the model generates far more tokens overall
  • Net recommendation: “Just don’t use Gemini.”


Sonnet 5 (low D / mostly F tier depending on environment)

  • Use case

    • Testing Claude Code sub-agents inside T3 Code
    • Visualizable math-ish tasks
    • Can spin up sub agents
  • But

    • Not token efficient
    • More expensive than it should be
  • Hard rule

    • Don’t pick it for anything except Cloud Code sub-agents
  • Tier: “Low D,” with strict usage constraints


Opus 5 (C / high D due to “mimic behavior”)

  • Initial impression

    • Seems capable of matching Fable-level detail
    • Catches things other models miss
  • Failure mode during merging

    • When code must actually run/merge:
      • outputs become vague/jargony
      • code execution quality is poor
  • Behavior critique

    • “Looks like a duck… mimic behavior”
    • It convinces you it’s smart until you try to integrate it
  • Tier

    • Low C capability but high D cost/behavior
    • Net ends up around high D

Anthropic Fable 5 (winner; S tier—top of chart)

  • Tier

    • S tier
    • Framed as “the only S tier model right now”
  • Strengths

    • Best overall code quality and depth of understanding
    • Most willing to merge output
    • Strong at:
      • double-checking work
      • handling deep/hard engineering exploration
  • How it differs from Soul (workflow + agentic behavior)

    • Soul is “thorough but writes worse code”
    • Fable understands goals better and produces code the speaker feels safer merging
    • With Fable/Soul, the speaker can give vague instructions and approve
    • Models then handle:
      • computer-use verification
      • sub-agent review before PR creation
  • Tradeoffs / imperfections

    • Trips over itself sometimes
    • Loses track occasionally
    • Takes unnecessary shortcuts
    • Touches things it shouldn’t
    • Some issues attributed to Anthropic-side constraints
  • Pricing/business note

    • Speaker says they would never pay full API price
    • Relies on sub discounts (they “burn” a lot of subscription value weekly)

Notable “guides/tips” embedded in the tier logic

  • Use Luna/Flash as background context processors rather than irreversible “merge code” agents.
  • Treat non-vision models as unacceptable if you frequently paste images (explicit demotion rule applied to V4 Flash and V4 Pro).
  • Choose reasoning tiers based on token efficiency:
    • higher reasoning can explode token counts and reduce cost-value
  • Prefer models that fit your agent workflow:
    • If you rely on iterative decomposition + sub-agent orchestration, some models work better
    • If you want “tell me what to build and it will PR,” the best models differ fundamentally

Sponsors / product feature mentioned (technology context)

  • Browserbase
    • Described as an agent-friendly browser solution that provides web access for tasks like:
      • searching like Google
      • fetching known URLs
      • full interactive browsing (including sign-in/click actions)
    • Used to argue that models without web context feel less capable
    • Browser access improves agent intelligence

Main speakers/sources

  • Main speaker: The solo narrator/reviewer (no specific name provided in the subtitles)
  • Primary sources/models discussed:
    • 56 Soul, 56 Terra, 56 Luna
    • DeepSeek V4 Flash/Pro
    • Kimmy K3
    • GLM 53
    • Composer 2.5 (Composer 2.x lineage)
    • Grock 46
    • Muse Spark
    • Gemini (various Flash/Pro variants)
    • Anthropic Sonnet 5
    • Anthropic Opus 5
    • Anthropic Fable 5
  • Sponsor source: Browserbase
  • Also mentioned: a hiring sponsor, G2I

Original video