Video summary

Самое крупное обновление GOOGLE

Main summary

Key takeaways

Technology

Scale of Google’s AI usage (tokens / compute)

  • Google’s AI “token” processing is described as a unit of account for AI workload.
  • In 2024, it was cited as ~10B tokens/month across platforms, while Google is currently processing about 3.2 quadrillion tokens/month.
  • Token consumption is described as growing roughly 7× per year.
  • Google has 13 products with 1B+ users each (examples mentioned: YouTube and Gemini).

Gemini rollout across Google products (product-wide integration)

  • Google plans to embed Gemini into “almost every” service, aiming to make it helpful “everywhere and in everything.”
  • Google Maps example: users can ask for practical tasks (e.g., where to buy a dress within a time window). Gemini interprets and uses relevant data, including payment/card-related details.
  • YouTube “ask-first” interface (beyond recommendations):
    • Instead of only searching, you ask questions, and YouTube selects the best video.
    • Claimed advantage: YouTube has transcripts, so the system can understand video topics in advance.
    • The UI is described as a chat/LLM-like experience, including:
      • a mini-transcript
      • a breakdown by key points
      • “Click the video inside the chat” to open it at that moment
    • Context/memory claim: YouTube can remember the ongoing conversation.

Availability / regions

  • The new YouTube + Gemini chat mode is described as rolling out this summer in the US.

Search upgrades

  • Housing search demo: users describe requirements; Google searches across social networks and rental sites, not just a single site.
  • Proactive reminders: the system can notify users when something new appears.
  • Plans to add interactive “graph/visual” style answers similar to an Anthropic-style feature (interactive, informative visuals vs plain text).
  • Example themes mentioned for interactive visuals: learning/science, including gravitational waves.

Google Docs / productivity via chat

  • Gemini in docs enables users to:
    • talk live while Gemini drafts documents
    • use data from:
      • Google Drive
      • Email
      • additional online sources
  • Gemini can rewrite in simpler language and add structure such as tables.
  • “Create without touching the keyboard” capability is expected this summer.

New and upgraded Gemini models / multimodal generation

  • “Gemini Omni” is described as generating “anything” from input data (compared to Gemini content-generation tooling).
  • Referenced multimodal/video/world tools:
    • Veo (video generation)
    • Nanobana (image generation/editing)
    • a model for virtual worlds (walk/fly)
  • Demo concept: prompt like “Make a clay animation explaining protein folding.”
  • Video editing: users can upload a video and ask to add/remove/change elements (analogous to image editing workflows).

Detecting AI-generated / edited content (Synthetic ID)

  • Google introduced/expanded “Sint ID” (“SYN ID”) to label/augment search results with content provenance metadata.
  • Example use:
    • identify whether a photo was edited via Nanobana or processed through Google Photos.
  • Search can then answer questions like “was it generated?” and add context (e.g., whether it circulated online).
  • The approach is suggested to become more common due to industry adoption (mentions Nvidia, Open, and “Elon Laps” as parties implementing similar IDs).

Model performance

  • Gemini Flash 3.5 is described as a fast model (“thinks less”) but with stronger benchmark performance than a slower 3.1 Pro.
  • It’s positioned as very fast, and a review is promised (creator notes upcoming content).

“Antigravity” / agent tooling and developer features

  • Antigravity 2.0 updates include:
    • voice input
    • improved agent dialogue/orchestration
    • splitting complex tasks into sub-tasks handled by multiple agents
    • a new Antigravity CLI (not to be confused with “GMI CLI”)
  • Integration mentioned with:
    • Android
    • Firebase
    • Google Studio
  • Additional reference: using the engine inside Google Search for dynamic behavior and code execution.

Autonomous agent: Gemini Spark

  • Gemini Spark is described as a personal AI agent running on dedicated Google servers.
  • It uses Antigravity and MCP to connect to third-party services (framed as “OpenAI-like only in Google’s cloud”).
  • Availability: only for AI Ultra subscribers.
  • Additional mentions:
    • Spark added to Chrome so it can browse/act for the user
    • A Universal Card / single basket concept: users collect items from YouTube / Gemini / mail into one shared basket that checks price and stock in the background
    • Spark/agent capabilities later integrated into the Gemini app and other Google surfaces

Gemini app UX and settings

  • Gemini interface redesign: “more attractive” with updated themes (a “black” style mentioned).
  • Voice mode: choosing different English dialects.
  • A daily “brief” feature: Gemini reviews links/data (email, calendar, task book) and reminds you what to do.

Mac app and multimodal editing tools

  • Gemini app for Mac: select multiple files and ask (via a keyboard function key or voice) to draft something like a letter; Spark integrated for actions.
  • A separate Google Pix/Photos tool:
    • hover/select parts of an image
    • delete/drag/add text
    • create banners/posters from photos

Flow updates (video generation/editing)

  • Flow can now:
    • upload a photo and ask it to choose best angles
    • return multiple video options (e.g., 16)
    • regenerate those options with additional prompts (e.g., “from morning to night” phrasing)
  • Flow Tools: apply presets to videos/photos or create from ready-made templates without typing prompts.
  • Music generation is mentioned as improving instruments/presets.

Audio-focused smart glasses

  • Wearable described as audio smart glasses (no reliance on text on a display; audio only).
  • Demo includes:
    • using Maps navigation
    • asking Gemini/Spark to do tasks in the background on Google servers (e.g., “Order some coffee”)
  • Claim: Nanobana-type image processing works via glasses as well (take photo → ask to process).
  • Collaboration mentioned with Samsung and cross-OS support (iOS and Android).
  • Multiple models expected “in the fall.”

Security and research AI tooling

  • Cybersecurity agent “Code…/Codenr” (name partially garbled):
    • a dedicated Google agent for protecting codebases
    • automatically finding and fixing critical vulnerabilities
  • “Jmy for Science” toolkit for accelerating research:
    • finding relevant preprints
    • converting research into working code
    • proposing hypotheses
    • running simulations
  • Science examples:
    • Alpha Earth / digital twin concepts for deforestation and food security
    • weather forecasting model predicting a hurricane path 3 days early
    • future plans for virtual cell simulations
    • Isomorphic Labs using AI to model molecular interactions to develop drugs faster, mentioned as working in preclinical cancer/immune disorder areas

Main speakers/sources (as referenced in the subtitles)

  • “Google” / Google presenters at the main Google presentation
  • Hasabis (DeepMind/Google leadership; referenced while speaking about security)
  • The video creator/host (narrator guiding viewers, referencing their Telegram channel and promising reviews)

Original video