Video summary

GOOGLE IS BACK! (Gemini 3.8 Flash)

Main summary

Key takeaways

Technology

Summary of technological concepts & product evaluation (Gemini 3.8 Flash)

Competitive context

The video frames Google’s progress as initially lagging—after early missed ground versus ChatGPT—then briefly peaking with Gemini 2.5 Pro. Over the last ~6 months, Gemini releases are described as “pretty good” on benchmarks, but less impressive in real-world usage.

Main product

Google Gemini 3.8 Flash (released shortly after Gemini 3.7 Flash) is presented as a meaningful improvement—especially for price-to-performance.


Benchmark-focused review (what each test measures)

  1. DeepSUI (long-horizon software engineering) — highest priority

    • Interpreted as a proxy for model effectiveness on long-horizon coding and tasks.
    • Gemini 3.8 Flash: 73.7%
    • GPT 5.6 “o”-style score: 72.7%
    • Claim: Gemini is very close to top competitors (Claude is also mentioned as near-equivalent).
  2. GPD “val” (knowledge work: extraction + analysis + presentation)

    • Measures more business/real-world workflows like PDF extraction, data analysis, and presentation creation.
    • Gemini 3.8 Flash: 1545
    • Claude is described as dominating (1824); GPT is second (1710).
    • Video conclusion: Gemini is “just okay” here, roughly comparable to another model tier (GPT 5.6 “Terra” referenced).
  3. Harvey Legal benchmark (legal agent performance)

    • Gemini 3.8 Flash: #1 with 61.4%
    • Video asserts Gemini 3.8 Flash is “the best model on the planet” for legal work, while also describing it as cheap.
  4. Terminal Bench (coding benchmarks)

    • Terminal Bench 2.1: Gemini 3.8 Flash #1 at 89.4%
    • Terminal Bench 4.0: Gemini scores 19.1% (described as much lower)
    • Explanation offered: 4.0 is newer and harder; earlier benchmarks may have been closer to “saturated,” explaining the earlier strength.
    • Claude “Opus 5” is said to dominate Terminal Bench 4.0.
  5. “Humanity’s last exam”

    • Gemini 3.8 Flash: #1 at 55.9%
    • Noted as a surprising strong result compared to weaker outcomes elsewhere.
  6. OSWorld (agentic computer use)

    • Evaluates how well an agent controls a computer/browser.
    • Gemini 3.8 Flash: 59%
    • Claude “Opus 5” is higher at 75%.

Price-to-performance analysis (key takeaway)

  • The video emphasizes a combined chart: DeepSUI performance plus “average cost per task.”
  • Why it matters: even if a model is cheaper per token, it may use more tokens, so the effective cost per completed task may not improve.
  • Claimed outcome: Gemini 3.8 Flash is positioned “high up and to the right” on the cost-adjusted DeepSUI chart—good and cost-effective.

Product/variant: Gemini 3.8 Flash Cyber (cyber capabilities)

What it is

A second model variant, Gemini Flash 3.8 Cyber, targeted for cyber defense/attack-related capabilities.

Access restriction

Only available to “trusted defenders” via a Fair Wind program; general users can’t test it freely.

Benchmarking highlights

  • CyberJim benchmark: Gemini 3.8 Flash Cyber scores 86.2%
  • Other cyber-focused models are mentioned, including GPT cyber at 85.6%.

More realistic internal evaluation (multi-language vulnerability discovery)

  • The video claims the evaluation goes beyond C/C++-only tests.
  • It spans ~20 programming languages, aiming to discover a wide range of vulnerabilities across complex codebases.
  • Reported claim: Gemini 3.8 Flash Cyber is a massive jump vs Gemini 3.7 Flash on this internal test.

Practical demos/tests shown (creative & application tasks)

  1. 3D low-poly biome generation (7 terrains)

    • Results described as impressive, with beach and farm looking strong.
    • Qualitative comparison:
      • Soul: best detail (per presenter)
      • Fable: second
      • Gemini 3.8 Flash: around same as GLM 5.3, with some glitches/mistakes
  2. Web page generation for product prompts

    • Products included: Apple, NVIDIA DJX Spark, Rubber Duck Company, Galaxy Z Fold, Tesla Model Y.
    • Overall conclusion: mixed results—some look decent, others are incorrect/weird.
    • Tesla demo described as Microsoft-Paint-like, though still containing interactive/parameter-like fields and animation.
  3. PowerPoint deck generation about data centers

    • The model reportedly:
      • Uses a specific brand/theme (“Forward Future” branding/colors)
      • Produces structurally correct slide content
    • Presenter suggests Claude may be better for design polish.
  4. 3D topographic map of Mount Everest (interactive mapping)

    • Described as phenomenal:
      • Draggable/zoomable 3D terrain
      • Controls like cut angle, vertical exaggeration, and other parameters
      • Shows point locations (e.g., camp 2) within the 3D view
  5. Simple “Doom-like” recreation

    • Presented as a lightweight interactive demo—decent fun with additional prompts.

Overall conclusions from the video

  • Gemini 3.8 Flash is “good” in its strong areas:
    • DeepSUI engineering performance
    • Harvey legal benchmark performance
    • Some exam-style evaluation
  • Weaker spots include:
    • Knowledge-work benchmarks like GPD/real-world knowledge work
    • Some coding variants such as Terminal Bench 4.0
  • Standout differentiator: cost—positioned as extremely competitive per completed task, not just per-token pricing.
  • Cyber variant: highly capable but gated through the defender program.

Main speakers / sources

Main speaker

  • The video author/reviewer (narrator) presents benchmarks and demos, including statements like “I tested…,” chart walkthroughs, and references to a team member Alex sending links/tests.

Sources referenced in benchmarks

  • OpenAI (for the “GPD val” knowledge-work benchmark)
  • Claude (via “Claude Opus 5” competitor results)
  • Harvey (legal benchmark)
  • Terminal Bench (coding benchmarks)
  • OSWorld, DeepSUI, CyberJim (evaluation suites)

Original video