Video summary
تازهترین اخبار هوش مصنوعی: مدلهای جدید مایکروسافت و انویدیا، و جی پی تی
Main summary
Key takeaways
Summary of the video’s main points (AI news & model updates)
Opening / context
- The presenter returns after a long gap and frames the episode as an AI news update.
- They express condolences and hope for better days, referencing recent tragic events (including the “Mina school”).
- They also note Iran’s internet was reconnected, which delayed earlier publishing.
Microsoft’s announcements (major focus)
7 new model releases
The presenter highlights 7 new Microsoft model releases, covering:
- language/thinking models
- coding
- image generation
- speech-to-text transcription
- audio generation
Positioning strategy
- The presenter argues Microsoft is aiming for reliable, capable, state-of-the-art “independent” performance, not necessarily the absolute strongest/newest model in every category—especially for coding.
Arena-style leaderboard notes
In an Arena-like comparison where users vote/compare quality:
- Microsoft image model is ranked D
- described as strong, “after GPT”.
- Speech transcription model is emphasized as:
- 2–3x faster
- relatively low error
- but the presenter claims it does not yet support Persian
- based on shown language support (English and others mentioned).
Microsoft “Solara” platform + agentic devices
Solara: beyond chatbots
- Microsoft introduces Solara, intended to enable agentic capabilities beyond chat—including real-world actions on devices.
New device concepts
- A table-top conversational device (Alexa-like) that can perform task actions such as:
- calendar operations
- file-related actions
- A badge-like device with a screen/camera that can scan/record conversations and images for workplace/industry use cases, such as:
- doctor scanning barcodes
- industrial scanning
Debate on hardware vs software
- The presenter notes there’s debate whether these capabilities require a dedicated device.
- They also say Microsoft is trialing market response.
NVIDIA + Microsoft collaboration: “RTX Spark”
- NVIDIA and Microsoft describe a high-end Windows PC/supercomputer concept called RTX Spark.
- Key technical claim:
- CPU and GPU are unified on one chip
- using Unified Memory
- targeting up to 128GB, similar to Apple’s approach
- Intended benefits:
- stronger performance for AI inference
- including image/video/game/editing workloads
Pricing: unknown
-
Speculation: ~$3,000–$4,000+
-
Microsoft also announces Surface Ultra using the same chip:
- expected to be high performance
- cost still unknown
Coding / leaderboard positioning (including a “Chinese market update”)
The presenter gives an overview of “best model” options for coding/app development:
- Cloud Opus 4.7 / 4.7 Thinking:
- presented as the top choice for coding + creativity/thinking
- Mentions strong open-source coding, e.g.:
- K / Alibaba’s Qwen 3.7
- Mentions mixed rankings:
- some Gemini/GPT variants are described as lower than Cloud in the presenter’s assessment
- DeepSeek:
- described as creating hype but viewed as not matching expectations
- Gemini variants:
- framed as good multi-purpose models
- especially for image analysis to generate outputs
- but less strong for coding than Cloud
More model releases & trends (open-source, smaller + faster)
-
Nemotron Strata (NVIDIA)
- oriented toward long-term agent work rather than everyday chat
- open-source
- very large (~550B parameters), likely requiring cloud/API rather than local use
-
Gemma 4 (Google)
- open-source
- “unified architecture” (presented as replacing older “multimodal” framing with a single architecture)
- emphasized as able to run on smaller hardware, with example needs like:
- laptop class: ~16GB RAM
- GPU class: ~16GB VRAM
- logic performance emphasized; smaller footprint claimed
-
MiniMax / Ems
- described as open-source coding models
- said to match strong proprietary options on many benchmarks
-
Composer inside Cursor IDE
- Cursor is framed as an agent-like coding environment with codebase awareness
- Composer 2.5 is described as:
- comparable coding accuracy
- lower cost
- especially versus Cursor subscription tiers that can spike under heavy usage
Image generation: new approach and key models
Structure-based generation (vs diffusion-style)
The presenter compares:
- diffusion-style image generation
- a newer structure-based method:
- models convert prompts into a structured concept hierarchy
- decide spatial placement of objects
- enable precise edits (change one element without disturbing everything else)
Model contenders
- GPT-Image / Google Nano-like models (referenced as contenders)
- A newer set of models described as:
- good quality
- improved control
- including strong consistent character generation behavior
Ideogram
- Ideogram
- emphasized as open-source
- strong prompt following
- claimed to edit photos well (add/remove elements)
- referenced as doing well in Arena comparisons (strong, though not necessarily #1)
Video generation and editing
-
Focus on a Google video model that:
- generates videos from prompts
- can edit existing video
- e.g., change gender/design
- reposition objects/people
- described as learning “world/physics-like” structure rather than only producing visually plausible output
-
The presenter highlights an example where the model generates video aligned to:
- provided graphics
- written instructions —presented as a more “reasoned/structured” workflow than earlier approaches.
Sound / voice models
- Mentions an audio/voice model (from “Meez” and “Lebs” in subtitles) that claims:
- open-source availability
- control over emotions and voice tone
- English-only limitation currently
- Notes sample code and reinforces the trend toward smaller open-source models that can run locally.
Google Gemini Flash + agentic work + app building
-
Gemini 3.5 Flash
- described as 2–3x faster than heavier models
- keeps similar accuracy for certain tasks
- useful for “identity work” / folder-based labeling examples
- e.g., naming photos based on content + task
-
Google AI Studio
- described as enabling Android app building inside the browser:
- prompts generate code + preview
- later connect to a phone for installation/testing
- described as enabling Android app building inside the browser:
- Similar iOS app creation concept mentioned via related tooling.
Cloud coding / “Codex”-style tooling
- Explains Cloud Code / Codex-type products:
- users don’t need deep coding knowledge
- the tool breaks tasks into steps and builds apps end-to-end
- Compared with earlier coding chat assistants that required more manual fixing:
- this approach aims to reduce “getting stuck”
- Mentions plugin/skill ecosystems and an example iOS workflow created through these systems.
Personal app anecdote
- The presenter says they built an iOS app called “Dreamweaver” using these tools:
- dream journaling
- dream interpretation inspired by Carl Jung psychology
- Claims AI coding accelerated:
- frontend/UI development
- backend interpretation logic partly handled by the presenter
Google’s inbox-focused AI + additional research
-
Highlights an idea: AI inside the inbox that can:
- read/summarize emails
- propose actions
- voice-read messages
- enable replying
-
Additional Google research/tools mentioned:
- “Notebook”-style summarization of sources (text/PDF/video) and generation of study assets (including podcasts/graphs)
- CoScientist (agent for scientific research):
- runs through literature
- generates hypotheses
- critiques/improves hypotheses
- example claims include:
- proposing a drug regimen for rare leukemia
- findings related to liver fibrosis
- and a vision-related problem
Presenters / contributors mentioned
- Microsoft (model/platform announcements)
- Google (announcements, Gemini models, AI Studio, research systems)
- NVIDIA (RTX Spark collaboration; Nemotron Strata)
- Open-source model developers/companies referenced
- Alibaba (Qwen 3.7)
- DeepSeek
- MiniMax
- Kimi
- Ideogram
- Cursor/Composer ecosystem
- Entropy (Codex/Cloud Code described)
- Carl Jung (referenced via “Dreamweaver” interpretation framework)
- The video’s presenter (speaker): not named in the subtitles (addressed as “friends” / “I am back…” only)