Video summary
Did OpenAI actually build AGI? GPT-6 Astra first look
Main summary
Key takeaways
Tech/news summary (GPT-6 Astra and competing model releases)
Anthropic (Tuesday): Fable & Mythos 5.1
- Presented as “advanced” coding + knowledge-work models.
- Two-model split by availability:
- Mythos: described as not usable
- Fable: described as the available one
- Case study—debugging crash:
- A hedge fund used Fable 5.1 to analyze a program crash by using a memory snapshot.
- They located a crash address inside a compiled vendor library.
- Then they reconstructed/disassembled it to trace the issue to a vendor bug (no source access).
- Biotech—protein target binding:
- Claimed improvement from ~10% to ~50% success in designing proteins that bind specific body targets.
- Space/geoscience—Venus mapping:
- Used 30-year-old NASA radar data to train a neural network to build an elevation map for a region of Venus.
Meta (Wednesday): Muse Spark 1.3
- Framed as a frontier model and the fourth release in ~5 months.
- Pricing:
- Standard endpoint: $1.25 in / $4.25 out
- Contributor tier:
- 10¢ in / 20¢ out, but requires allowing Meta to train on user submissions (positioned as attractive due to cost tradeoffs).
OpenAI (Thursday): GPT-6 Astra (“AGI” claim)
- Video focuses on whether OpenAI “actually built AGI,” using an “AGI if you got access” framing for selected influencers/enterprise.
- Release-day disruption:
- An auto-generated claim said ChatGPT, Claude, Grok, and Cursor went down around the time of Astra coverage.
- Speculation ranged from a likely Azure outage to a more conspiratorial “kill the competition” theory.
- Launch/embargo chaos:
- OpenAI launched a page; outlets published embargoed stories.
- OpenAI allegedly took down the announcement.
- Influencers competed on how long they’d had secret access.
- About 90 minutes later, Astra was said to not yet be public; rollout to Plus/Pro was expected over the next days.
- Sam Altman posted an apology; a “go to bed” response is quoted when asked whether to stay up for availability.
- Claim: the model underwent a formal review with the Trump administration prior to release.
Astra technical/product claims
Training scale / infrastructure
- Pre-trained on 100,000+ GPUs at Stargate (Texas).
- Claimed distinction vs older OpenAI models:
- Earlier work involved more supervising during training
- Astra presumably shifts toward a different approach.
Core capability marketed: “computer use”
- Can fill forms, crunch spreadsheets, and operate engineering tools including KiCad and Blender.
Desktop agent benchmark results
- OSWorld benchmark:
- 73% score
- ~40 minutes per task
- Compared to Soul: 65% but ~75 minutes
- Trust Me Bro benchmark:
- Claims: 100% on exploit bench, 65% on terminal bench
- Other notable results are contrasted with earlier Anthropic “bragging” metrics.
- Arc AGI 3:
- 99%
- Presented as evidence of generalization rather than memorization
- Tied to discussion around “AGI” definition debates.
Cyber preparedness threshold
- OpenAI’s claim: Astra reached a “critical cyber threshold” in a preparedness framework.
- Framed capability:
- Able to find and exploit zero-days without a human specifying actions
- Video treats this as a major differentiator.
Pricing
- Astra pricing matched to Fable 5.1:
- $10 per 1M input tokens / $50 per 1M output tokens
Early access reviews / demos highlighted
Overall tone
- Described as positive early-access reviews, typical of access groups (not loudly critical publicly).
Spatial/world understanding demos
- Sherif Shamim:
- Recreated the Palace of Fine Arts (San Francisco) in Blender, described as nearly perfect.
- OpenAI staff (Thomas Recart):
- Walked through a demo house—modeled in Blender and converted into a fully walkable Unreal Engine 5 scene.
- Matt Schumer:
- Created a world in Unreal Engine populated with dozen-ish Astra agents.
- Reported immersive behavior (agents in his home “talking” about plans—comedic interpretation in the video).
Independent evaluation discrepancy
- Artificial Analysis – Independent Intelligence Index:
- Astra scored 61
- Described as exactly equal to GPT-5.6 Soul
- And 5 points behind Anthropic’s Fable 5.1
- The video suggests the results don’t align with Astra’s strong benchmark/cyber claims.
Sponsor / security tool mention (Code Rabbit)
- Code Rabbit Security (sponsor) is mentioned as providing “continuous code security.”
- Positioning:
- Uses reasoning-based agents rather than regex/brittle rules
- Agents are intended to think like an attacker and scan for vulnerabilities across a codebase
- Prioritization factors:
- reachability, exploitability, blast radius
- Outputs:
- plain-English explanations
- a suggested fix (approve/merge from a diff)
- Claims:
- Reviews each PR before merge
- Supports scheduled deep scans
- Offer:
- 10 free code scans
Main speakers/sources (as referenced in the subtitles)
- Narrator/host: The Code Report (speaker throughout)
- Sam Altman: quoted on rollout/apology and claims about government review
- President Greg: quoted on Astra access/AGI framing
- Alexander Wang: quoted about Meta’s contributor tier adoption
- Sherif Shamim: spatial demo creator
- Thomas Recart (OpenAI): demo walkthrough author
- Matt Schumer: Unreal Engine agent demo
- Artificial Analysis: publisher of the Independent Intelligence Index score
- Code Rabbit: sponsor (presented via host endorsement)