Video summary

Did OpenAI actually build AGI? GPT-6 Astra first look

Main summary

Key takeaways

Technology

Tech/news summary (GPT-6 Astra and competing model releases)

Anthropic (Tuesday): Fable & Mythos 5.1

  • Presented as “advanced” coding + knowledge-work models.
  • Two-model split by availability:
    • Mythos: described as not usable
    • Fable: described as the available one
  • Case study—debugging crash:
    • A hedge fund used Fable 5.1 to analyze a program crash by using a memory snapshot.
    • They located a crash address inside a compiled vendor library.
    • Then they reconstructed/disassembled it to trace the issue to a vendor bug (no source access).
  • Biotech—protein target binding:
    • Claimed improvement from ~10% to ~50% success in designing proteins that bind specific body targets.
  • Space/geoscience—Venus mapping:
    • Used 30-year-old NASA radar data to train a neural network to build an elevation map for a region of Venus.

Meta (Wednesday): Muse Spark 1.3

  • Framed as a frontier model and the fourth release in ~5 months.
  • Pricing:
    • Standard endpoint: $1.25 in / $4.25 out
  • Contributor tier:
    • 10¢ in / 20¢ out, but requires allowing Meta to train on user submissions (positioned as attractive due to cost tradeoffs).

OpenAI (Thursday): GPT-6 Astra (“AGI” claim)

  • Video focuses on whether OpenAI “actually built AGI,” using an “AGI if you got access” framing for selected influencers/enterprise.
  • Release-day disruption:
    • An auto-generated claim said ChatGPT, Claude, Grok, and Cursor went down around the time of Astra coverage.
    • Speculation ranged from a likely Azure outage to a more conspiratorial “kill the competition” theory.
  • Launch/embargo chaos:
    • OpenAI launched a page; outlets published embargoed stories.
    • OpenAI allegedly took down the announcement.
    • Influencers competed on how long they’d had secret access.
    • About 90 minutes later, Astra was said to not yet be public; rollout to Plus/Pro was expected over the next days.
    • Sam Altman posted an apology; a “go to bed” response is quoted when asked whether to stay up for availability.
    • Claim: the model underwent a formal review with the Trump administration prior to release.

Astra technical/product claims

Training scale / infrastructure

  • Pre-trained on 100,000+ GPUs at Stargate (Texas).
  • Claimed distinction vs older OpenAI models:
    • Earlier work involved more supervising during training
    • Astra presumably shifts toward a different approach.

Core capability marketed: “computer use”

  • Can fill forms, crunch spreadsheets, and operate engineering tools including KiCad and Blender.

Desktop agent benchmark results

  • OSWorld benchmark:
    • 73% score
    • ~40 minutes per task
    • Compared to Soul: 65% but ~75 minutes
  • Trust Me Bro benchmark:
    • Claims: 100% on exploit bench, 65% on terminal bench
    • Other notable results are contrasted with earlier Anthropic “bragging” metrics.
  • Arc AGI 3:
    • 99%
    • Presented as evidence of generalization rather than memorization
    • Tied to discussion around “AGI” definition debates.

Cyber preparedness threshold

  • OpenAI’s claim: Astra reached a “critical cyber threshold” in a preparedness framework.
  • Framed capability:
    • Able to find and exploit zero-days without a human specifying actions
    • Video treats this as a major differentiator.

Pricing

  • Astra pricing matched to Fable 5.1:
    • $10 per 1M input tokens / $50 per 1M output tokens

Early access reviews / demos highlighted

Overall tone

  • Described as positive early-access reviews, typical of access groups (not loudly critical publicly).

Spatial/world understanding demos

  • Sherif Shamim:
    • Recreated the Palace of Fine Arts (San Francisco) in Blender, described as nearly perfect.
  • OpenAI staff (Thomas Recart):
    • Walked through a demo house—modeled in Blender and converted into a fully walkable Unreal Engine 5 scene.
  • Matt Schumer:
    • Created a world in Unreal Engine populated with dozen-ish Astra agents.
    • Reported immersive behavior (agents in his home “talking” about plans—comedic interpretation in the video).

Independent evaluation discrepancy

  • Artificial Analysis – Independent Intelligence Index:
    • Astra scored 61
    • Described as exactly equal to GPT-5.6 Soul
    • And 5 points behind Anthropic’s Fable 5.1
  • The video suggests the results don’t align with Astra’s strong benchmark/cyber claims.

Sponsor / security tool mention (Code Rabbit)

  • Code Rabbit Security (sponsor) is mentioned as providing “continuous code security.”
  • Positioning:
    • Uses reasoning-based agents rather than regex/brittle rules
    • Agents are intended to think like an attacker and scan for vulnerabilities across a codebase
  • Prioritization factors:
    • reachability, exploitability, blast radius
  • Outputs:
    • plain-English explanations
    • a suggested fix (approve/merge from a diff)
  • Claims:
    • Reviews each PR before merge
    • Supports scheduled deep scans
  • Offer:
    • 10 free code scans

Main speakers/sources (as referenced in the subtitles)

  • Narrator/host: The Code Report (speaker throughout)
  • Sam Altman: quoted on rollout/apology and claims about government review
  • President Greg: quoted on Astra access/AGI framing
  • Alexander Wang: quoted about Meta’s contributor tier adoption
  • Sherif Shamim: spatial demo creator
  • Thomas Recart (OpenAI): demo walkthrough author
  • Matt Schumer: Unreal Engine agent demo
  • Artificial Analysis: publisher of the Independent Intelligence Index score
  • Code Rabbit: sponsor (presented via host endorsement)

Original video