Video summary
Самое крупное обновление GOOGLE
Main summary
Key takeaways
Scale of Google’s AI usage (tokens / compute)
- Google’s AI “token” processing is described as a unit of account for AI workload.
- In 2024, it was cited as ~10B tokens/month across platforms, while Google is currently processing about 3.2 quadrillion tokens/month.
- Token consumption is described as growing roughly 7× per year.
- Google has 13 products with 1B+ users each (examples mentioned: YouTube and Gemini).
Gemini rollout across Google products (product-wide integration)
- Google plans to embed Gemini into “almost every” service, aiming to make it helpful “everywhere and in everything.”
- Google Maps example: users can ask for practical tasks (e.g., where to buy a dress within a time window). Gemini interprets and uses relevant data, including payment/card-related details.
- YouTube “ask-first” interface (beyond recommendations):
- Instead of only searching, you ask questions, and YouTube selects the best video.
- Claimed advantage: YouTube has transcripts, so the system can understand video topics in advance.
- The UI is described as a chat/LLM-like experience, including:
- a mini-transcript
- a breakdown by key points
- “Click the video inside the chat” to open it at that moment
- Context/memory claim: YouTube can remember the ongoing conversation.
Availability / regions
- The new YouTube + Gemini chat mode is described as rolling out this summer in the US.
Search upgrades
- Housing search demo: users describe requirements; Google searches across social networks and rental sites, not just a single site.
- Proactive reminders: the system can notify users when something new appears.
- Plans to add interactive “graph/visual” style answers similar to an Anthropic-style feature (interactive, informative visuals vs plain text).
- Example themes mentioned for interactive visuals: learning/science, including gravitational waves.
Google Docs / productivity via chat
- Gemini in docs enables users to:
- talk live while Gemini drafts documents
- use data from:
- Google Drive
- additional online sources
- Gemini can rewrite in simpler language and add structure such as tables.
- “Create without touching the keyboard” capability is expected this summer.
New and upgraded Gemini models / multimodal generation
- “Gemini Omni” is described as generating “anything” from input data (compared to Gemini content-generation tooling).
- Referenced multimodal/video/world tools:
- Veo (video generation)
- Nanobana (image generation/editing)
- a model for virtual worlds (walk/fly)
- Demo concept: prompt like “Make a clay animation explaining protein folding.”
- Video editing: users can upload a video and ask to add/remove/change elements (analogous to image editing workflows).
Detecting AI-generated / edited content (Synthetic ID)
- Google introduced/expanded “Sint ID” (“SYN ID”) to label/augment search results with content provenance metadata.
- Example use:
- identify whether a photo was edited via Nanobana or processed through Google Photos.
- Search can then answer questions like “was it generated?” and add context (e.g., whether it circulated online).
- The approach is suggested to become more common due to industry adoption (mentions Nvidia, Open, and “Elon Laps” as parties implementing similar IDs).
Model performance
- Gemini Flash 3.5 is described as a fast model (“thinks less”) but with stronger benchmark performance than a slower 3.1 Pro.
- It’s positioned as very fast, and a review is promised (creator notes upcoming content).
“Antigravity” / agent tooling and developer features
- Antigravity 2.0 updates include:
- voice input
- improved agent dialogue/orchestration
- splitting complex tasks into sub-tasks handled by multiple agents
- a new Antigravity CLI (not to be confused with “GMI CLI”)
- Integration mentioned with:
- Android
- Firebase
- Google Studio
- Additional reference: using the engine inside Google Search for dynamic behavior and code execution.
Autonomous agent: Gemini Spark
- Gemini Spark is described as a personal AI agent running on dedicated Google servers.
- It uses Antigravity and MCP to connect to third-party services (framed as “OpenAI-like only in Google’s cloud”).
- Availability: only for AI Ultra subscribers.
- Additional mentions:
- Spark added to Chrome so it can browse/act for the user
- A Universal Card / single basket concept: users collect items from YouTube / Gemini / mail into one shared basket that checks price and stock in the background
- Spark/agent capabilities later integrated into the Gemini app and other Google surfaces
Gemini app UX and settings
- Gemini interface redesign: “more attractive” with updated themes (a “black” style mentioned).
- Voice mode: choosing different English dialects.
- A daily “brief” feature: Gemini reviews links/data (email, calendar, task book) and reminds you what to do.
Mac app and multimodal editing tools
- Gemini app for Mac: select multiple files and ask (via a keyboard function key or voice) to draft something like a letter; Spark integrated for actions.
- A separate Google Pix/Photos tool:
- hover/select parts of an image
- delete/drag/add text
- create banners/posters from photos
Flow updates (video generation/editing)
- Flow can now:
- upload a photo and ask it to choose best angles
- return multiple video options (e.g., 16)
- regenerate those options with additional prompts (e.g., “from morning to night” phrasing)
- Flow Tools: apply presets to videos/photos or create from ready-made templates without typing prompts.
- Music generation is mentioned as improving instruments/presets.
Audio-focused smart glasses
- Wearable described as audio smart glasses (no reliance on text on a display; audio only).
- Demo includes:
- using Maps navigation
- asking Gemini/Spark to do tasks in the background on Google servers (e.g., “Order some coffee”)
- Claim: Nanobana-type image processing works via glasses as well (take photo → ask to process).
- Collaboration mentioned with Samsung and cross-OS support (iOS and Android).
- Multiple models expected “in the fall.”
Security and research AI tooling
- Cybersecurity agent “Code…/Codenr” (name partially garbled):
- a dedicated Google agent for protecting codebases
- automatically finding and fixing critical vulnerabilities
- “Jmy for Science” toolkit for accelerating research:
- finding relevant preprints
- converting research into working code
- proposing hypotheses
- running simulations
- Science examples:
- Alpha Earth / digital twin concepts for deforestation and food security
- weather forecasting model predicting a hurricane path 3 days early
- future plans for virtual cell simulations
- Isomorphic Labs using AI to model molecular interactions to develop drugs faster, mentioned as working in preclinical cancer/immune disorder areas
Main speakers/sources (as referenced in the subtitles)
- “Google” / Google presenters at the main Google presentation
- Hasabis (DeepMind/Google leadership; referenced while speaking about security)
- The video creator/host (narrator guiding viewers, referencing their Telegram channel and promising reviews)