Video summary
GOOGLE IS BACK! (Gemini 3.8 Flash)
Main summary
Key takeaways
Summary of technological concepts & product evaluation (Gemini 3.8 Flash)
Competitive context
The video frames Google’s progress as initially lagging—after early missed ground versus ChatGPT—then briefly peaking with Gemini 2.5 Pro. Over the last ~6 months, Gemini releases are described as “pretty good” on benchmarks, but less impressive in real-world usage.
Main product
Google Gemini 3.8 Flash (released shortly after Gemini 3.7 Flash) is presented as a meaningful improvement—especially for price-to-performance.
Benchmark-focused review (what each test measures)
-
DeepSUI (long-horizon software engineering) — highest priority
- Interpreted as a proxy for model effectiveness on long-horizon coding and tasks.
- Gemini 3.8 Flash: 73.7%
- GPT 5.6 “o”-style score: 72.7%
- Claim: Gemini is very close to top competitors (Claude is also mentioned as near-equivalent).
-
GPD “val” (knowledge work: extraction + analysis + presentation)
- Measures more business/real-world workflows like PDF extraction, data analysis, and presentation creation.
- Gemini 3.8 Flash: 1545
- Claude is described as dominating (1824); GPT is second (1710).
- Video conclusion: Gemini is “just okay” here, roughly comparable to another model tier (GPT 5.6 “Terra” referenced).
-
Harvey Legal benchmark (legal agent performance)
- Gemini 3.8 Flash: #1 with 61.4%
- Video asserts Gemini 3.8 Flash is “the best model on the planet” for legal work, while also describing it as cheap.
-
Terminal Bench (coding benchmarks)
- Terminal Bench 2.1: Gemini 3.8 Flash #1 at 89.4%
- Terminal Bench 4.0: Gemini scores 19.1% (described as much lower)
- Explanation offered: 4.0 is newer and harder; earlier benchmarks may have been closer to “saturated,” explaining the earlier strength.
- Claude “Opus 5” is said to dominate Terminal Bench 4.0.
-
“Humanity’s last exam”
- Gemini 3.8 Flash: #1 at 55.9%
- Noted as a surprising strong result compared to weaker outcomes elsewhere.
-
OSWorld (agentic computer use)
- Evaluates how well an agent controls a computer/browser.
- Gemini 3.8 Flash: 59%
- Claude “Opus 5” is higher at 75%.
Price-to-performance analysis (key takeaway)
- The video emphasizes a combined chart: DeepSUI performance plus “average cost per task.”
- Why it matters: even if a model is cheaper per token, it may use more tokens, so the effective cost per completed task may not improve.
- Claimed outcome: Gemini 3.8 Flash is positioned “high up and to the right” on the cost-adjusted DeepSUI chart—good and cost-effective.
Product/variant: Gemini 3.8 Flash Cyber (cyber capabilities)
What it is
A second model variant, Gemini Flash 3.8 Cyber, targeted for cyber defense/attack-related capabilities.
Access restriction
Only available to “trusted defenders” via a Fair Wind program; general users can’t test it freely.
Benchmarking highlights
- CyberJim benchmark: Gemini 3.8 Flash Cyber scores 86.2%
- Other cyber-focused models are mentioned, including GPT cyber at 85.6%.
More realistic internal evaluation (multi-language vulnerability discovery)
- The video claims the evaluation goes beyond C/C++-only tests.
- It spans ~20 programming languages, aiming to discover a wide range of vulnerabilities across complex codebases.
- Reported claim: Gemini 3.8 Flash Cyber is a massive jump vs Gemini 3.7 Flash on this internal test.
Practical demos/tests shown (creative & application tasks)
-
3D low-poly biome generation (7 terrains)
- Results described as impressive, with beach and farm looking strong.
- Qualitative comparison:
- Soul: best detail (per presenter)
- Fable: second
- Gemini 3.8 Flash: around same as GLM 5.3, with some glitches/mistakes
-
Web page generation for product prompts
- Products included: Apple, NVIDIA DJX Spark, Rubber Duck Company, Galaxy Z Fold, Tesla Model Y.
- Overall conclusion: mixed results—some look decent, others are incorrect/weird.
- Tesla demo described as Microsoft-Paint-like, though still containing interactive/parameter-like fields and animation.
-
PowerPoint deck generation about data centers
- The model reportedly:
- Uses a specific brand/theme (“Forward Future” branding/colors)
- Produces structurally correct slide content
- Presenter suggests Claude may be better for design polish.
- The model reportedly:
-
3D topographic map of Mount Everest (interactive mapping)
- Described as phenomenal:
- Draggable/zoomable 3D terrain
- Controls like cut angle, vertical exaggeration, and other parameters
- Shows point locations (e.g., camp 2) within the 3D view
- Described as phenomenal:
-
Simple “Doom-like” recreation
- Presented as a lightweight interactive demo—decent fun with additional prompts.
Overall conclusions from the video
- Gemini 3.8 Flash is “good” in its strong areas:
- DeepSUI engineering performance
- Harvey legal benchmark performance
- Some exam-style evaluation
- Weaker spots include:
- Knowledge-work benchmarks like GPD/real-world knowledge work
- Some coding variants such as Terminal Bench 4.0
- Standout differentiator: cost—positioned as extremely competitive per completed task, not just per-token pricing.
- Cyber variant: highly capable but gated through the defender program.
Main speakers / sources
Main speaker
- The video author/reviewer (narrator) presents benchmarks and demos, including statements like “I tested…,” chart walkthroughs, and references to a team member Alex sending links/tests.
Sources referenced in benchmarks
- OpenAI (for the “GPD val” knowledge-work benchmark)
- Claude (via “Claude Opus 5” competitor results)
- Harvey (legal benchmark)
- Terminal Bench (coding benchmarks)
- OSWorld, DeepSUI, CyberJim (evaluation suites)