Video summary
Why Opus 4.8 Pulled Me Back to Claude
Main summary
Key takeaways
Tech/Product Takeaways (What This Video Says About Opus 4.8)
Release context
- Opus 4.8 is framed as a major Anthropic jump—it’s suggested it “could have been Opus 5” because it feels meaningfully better than Opus 4.7.
Overall positioning
- The speaker claims Opus 4.8 is “top of the pack” on their internal benchmarks.
- In some areas, results are described as being close to GPT 5.5, and clearly better than 4.7.
Core strengths
-
Writing
- Rated as the best writing model they’ve tested.
- Writing benchmark score: 79.6/100 (vs GPT 5.5: 73).
- Behavior notes: Fast and expressive, with fewer “AI tells”, especially on high reasoning tasks.
- Voice imitation: Strong at continuing a user’s writing voice from context (e.g., “Continue” + your paragraph keeps your style).
- Voice/persona work: Especially good for personal/interpersonal and emotionally intelligent writing; described as “best at pushing my own frame.”
-
Knowledge work
- Strong at producing structured artifacts like slide decks.
- Example: it generated a beginner slide deck with depth and good styling—the first time the presenter felt an auto-generated deck had “real” richness.
- Also described as good for mixed threads (e.g., switching between coding + writing in the same conversation).
-
Coding
- A powerhouse on extra-high reasoning settings.
- Senior engineer benchmark: 63
- vs Opus 4.7: about 30 points lower
- vs GPT 5.5: 62
- Benchmark method (as described):
- The model gets a “vibe-coded slop” codebase.
- It must rewrite from first principles.
- Output is compared to rewrites by human senior engineers.
- LFG bench (more realistic style coding tasks):
- Examples include SaaS/e-commerce and 3D game landscapes.
- Output quality: code is readable and the model balances engineering quality with creative detail.
- Example comparison: a 3D “cozy island” scene is described as richer/lush on 4.8, while GPT 5.5 is more diverse but less “vibrant,” more straightforward.
Main caveats
-
Reasoning sensitivity
- Performance improves substantially with reasoning settings.
- Best results are emphasized for high and extra high (especially for hardest coding + important writing).
- Medium is noted as weaker, particularly for writing.
-
Daily-driver limitation: UI/product ecosystem
- Despite the model strength, the presenter prefers Codex’s desktop app over Claude’s desktop app, due to perceived UI/UX fragmentation.
- Specific complaint:
- Claude’s app has multiple tabs (chat/code/cowork) that feel like they’re run by different teams (“shipping an org chart”).
- Codex is described as faster, simpler, and it includes an in-app browser, which helps knowledge work.
-
Practical recommendation
- Don’t rely on Claude as your only interface yet.
- Use Opus 4.8 as part of your toolkit/arsenal.
- Expect best results using Claude desktop/code with high/extra-high reasoning.
Review / Benchmark Framework Mentioned
- “Day zero vibe check” and internal testing at Every for about a week.
- “Reach test”: a simple measure of whether you naturally want to use it (“do you reach for it, and when?”)
- Speaker: Gold / green (limited by harness/daily-driver UX)
- Kieran Klassen (GM of Quora): Straight Gold (“paradigm shift”)
- Katie Parrot (senior staff writer): Green (mostly writing/knowledge work)
- Senior Engineer benchmark
- Rewrite from first principles; compared against human senior engineers.
- LFG bench
- Real-world style coding tasks (e.g., SaaS/e-commerce, 3D environments).
- Writing benchmark
- Multiple writing genres (e.g., intro, promo email, middle paragraphs).
- Knowledge-work test
- Specifically slide deck generation quality and depth.
Tutorials / Guides Explicitly Referenced
- No step-by-step tutorial for Opus itself, but the speaker provides testing guidance:
- Use High / Extra High reasoning for the hardest coding and most important writing.
- For best experience, try it in Claude desktop app and Claude code.
- Compare workflow against Codex desktop app.
Main Speakers / Sources
- Speaker: Narrator from Every (host/tester; internal voices referenced).
- Mentioned internal testers/sources:
- Kieran Klassen — GM of Quora (internal tester; “most human model” comment; “Paradigm shift” grade)
- Katie Parrot — senior staff writer (internal tester; “reach test” rating)
- Primary product/company context: Anthropic (Claude Opus models).