Video summary

Did Elon catch up? (Grok 4.7 is here)

Main summary

Key takeaways

Product Review

Product reviewed

Grok 4.7 (xAI) — positioned as an upgrade to Grok 4.6, claiming improved reasoning/coding performance while keeping the same price.


Key features / what’s improved (from the video)

  • Coding + knowledge work focus

    • Framed as the “most capable model for coding and knowledge work.”
  • Longer on difficult tasks

    • Tends to work longer on harder problems rather than giving up early.
  • Self-checking / more careful work

    • “Checks its own work more carefully.”
  • Safeguards

    • “Best calibrated safeguards to date.”
    • Highlighted for strongest tested performance on:
      • Refusals
      • Jailbreak resistance
    • Notes improvements in dual-use domains, such as:
      • Cybersecurity
      • Biological work
    • Introduces a new safeguard stack, described as a major safety differentiator.

Pricing / cost claims

  • Same price as Grok 4.6, presented as very cost-effective.
  • Cost per tokens (given)
    • $2 per million input tokens
    • $6 per million output tokens
  • Comparisons
    • GPT 5.6: “more than twice as expensive”
    • Fable 5.1: “over five times more expensive”
    • GPT/Astra: about ~5x price for only a “few percentage point” improvement (per narrator)

Core argument: Many industries want automation/knowledge work, not absolute frontier performance—so Grok 4.7’s cost/performance is positioned as the main win.


Benchmark highlights & comparisons mentioned

CursorBench 4.0 (performance vs cost)

  • Setup: top-right is best (high performance, low cost).
  • Grok 4.7: “quite good.”

Notable comparisons

  • Grok 4.7 vs Opus 5 / GPT 5.x
    • Comparable, but not beating Opus 5 on max thinking.
    • Much less expensive (about half the cost to run the benchmark at that setting).
  • Fable 5.1
    • “Winner” on score, but extremely expensive, so cost per task completed becomes worse.

Tokens / “intelligence density” discussion

  • Another view of CursorBench uses average output tokens per task (lower is better).
  • Grok 4.7 is described as very comparable to Opus 5 in efficiency/usage patterns.

DeepSeek / “Deep Suite” coding benchmark (real-world coding feel)

Reported results

  • Grok 4.7: 71%
  • GPT 5.6: 72.7%
  • Fable 5.1: 70%

Narrator notes possible cherry-picking, mentioning Astra wasn’t shown in the initial table, then presents Astra separately.

Deep Suite including Astra (recreated table)

  • Astra Max: 74.1% (best)
  • GPT 5.6 Soul: 72.7%
  • Grok 4.7: 71%
  • Fable 5.1: 70%

Conclusion: Astra has best coding on Deep Suite, but Grok remains very strong.

Other benchmarks

  • Briefcase (office work)

    • Described as very competitive with Grok 4.7 and Fable 5.1.
  • Terminal Bench 4.0

    • Grok 4.7: 38%
    • Fable 5.1: 57.9%
    • Astra: 58.2%
    • Takeaway: Grok lags on terminal/agentic coding speed.
  • Legal Work

    • Claims Grok 4.7 dominated:
      • Grok 4.7: 19.6%
      • Grok 4.6: 15.8%
    • Other results shown as very low (e.g., GPT 5.6 Soul 2.5%, Fable 6.7 ~?, Astra ~5.4).
    • Also highlights Muse Spark 1.2 as “winner” for legal benchmark by the speaker.
    • Net: there’s some inconsistency in how “dominance” is framed, but Muse Spark is clearly emphasized for legal work in at least one narrative.
  • Healthbench Professional

    • “About the same across the board.”
  • Electrical Engineering Bench

    • “#2 with Astra included,” and would win if Astra were removed.

OpenAI real-world knowledge tasks benchmark (“ELO” style)

  • Fable 5.1: 1735 ELO (top)
  • Grok 4.7: 1695 ELO
  • Astra Max (mentioned): 1542 ELO (4th place)

Takeaway: Grok 4.7 performs near the top, but not #1.

Artificial Analysis Intelligence Index

  • Fable 5.1: #1
  • Astra: #2 (scores mentioned as the same as #1)
  • Claude Opus 5: #3
  • Muse Spark: #4 (open source mentioned)
  • Grok 4.7 extra high: #5
    • Grok 4.7 extra high shown at 46 (approx. placement)

Also claimed:

  • Grok 4.7 may not appear on a cost-per-intelligence index graph due to a display/selection issue, though Grok 4.6 indicates it’s inexpensive.

User experience / practical notes (from the video)

  • Grok 4.7 is described as “baking” into better performance, suggesting earlier versions may have been prematurely penalized.
  • Mentions a delay (“needs a few more days to cook”), implying early iteration issues—possibly including:
    • Penalization in response length
    • Lack of rigor in checking work

Pros (as stated)

  • Strong cost-effectiveness
    • Described as “incredibly cost effective” and “best calibrated safeguards.”
  • Better hard-task behavior
    • Stays with difficult tasks longer.
  • Improved safety
    • Strong refusals and jailbreak resistance.
  • Very competitive on multiple knowledge-work benchmarks
    • Including briefcase/office work and general knowledge tasks.
  • Good near-frontier performance at lower prices
    • Wins on cost per completed task in many comparisons.

Cons / limitations mentioned

  • Not always top performer
    • Example: Terminal Bench is significantly lower than Fable/Astra.
  • Smaller context window
    • Grok 4.6/4.7: 500K context
    • Competitors: “frontier” models around ~1M tokens
    • Impact mainly for specialized high-context tasks.
  • Benchmark cherry-picking concerns
    • Narrator suggests some benchmarks present Grok favorably and may omit Astra in at least one table until recreated.

Comparisons summary (who it’s compared against)

  • Opus 5 / Claude
    • Grok is comparable; not beating it at max thinking.
  • Fable 5.1
    • Often highest raw scores; Grok wins on cost per completed task.
  • GPT 5.6 / “Soul”
    • Sometimes slightly higher on specific benchmarks.
  • Astra Max
    • Sometimes best on Deep Suite and other tasks, but Grok remains strong on cost/performance.
  • Muse Spark
    • Highlighted for legal-work performance and open-source availability.

Named speaker/host contributions (as presented)

  • Primary narrator/reviewer

    • Provides benchmark commentary, pricing analysis, and overall recommendation.
  • Mentions of others/demos

    • Elon Musk tweets
      • Predicts Grok 4.7 would be “roughly on par with Opus 50,” plus later delay reasoning.
    • Cursorbench sponsor note
      • Automated benchmark references; user experience tied to the Cursor ecosystem.
    • “Bobby” demo
      • Grok 4.7 vs Kimi K3, with a harsh reaction (“terrible”), though narrator questions the settings used.
  • Sponsor: Zapier

    • Used to demonstrate integrations/workflows (not treated as a product critique), including claimed Grok 4.7 integration/automation use cases.

Overall verdict / recommendation

Recommendation: Buy/use Grok 4.7 if your priority is cost-effective coding + knowledge work with strong safety.

It’s described as near-frontier in many areas and significantly cheaper than top alternatives—often winning on cost per completed task. However, it’s not best-in-class across every benchmark (notably Terminal Bench), and it has a smaller 500K context window than many competitors.

Original video