Video summary
GPT-5.5, Claude 4.7, Gemini: Хватит платить за все нейронки!
Main summary
Key takeaways
Product reviewed (video’s subject)
The video is not a single product review; it’s a guide to which AI model to use for which task in 2026 (e.g., GPT-5.5, Claude 4.7/4.8, Gemini 3.1 Pro, Grok, DeepSeek, plus media models like GPT Image 2, Gemini/Google image & video tools, Runway, etc.).
Key message / overall recommendation
Don’t try to find “the best model.” In 2026, models are described as highly specialized, so the money-saving strategy is to use a small stack:
- 1 paid “base” model for ~80% of tasks (often Claude for tech, or Gemini for Google users)
- Free/cheap add-ons for specific needs (e.g., Grok for real-time X news, DeepSeek for bulk code under privacy)
This combination is claimed to save ~$2,000/year vs relying on a single universal model.
Unique points mentioned (by model/tool)
1) GPT Chat (GPT-5.5, released Apr 23, 2026)
Positioning / features
- Still described as the most versatile model; “not the best everywhere, but not the worst.”
- Has:
- Best overall intelligence rank (cited)
- 58.6 on SVE bench (code bug/applied tasks); compared to “IQ level”
- Strong plugin ecosystem + custom GPT
- Canvas editor for editing long documents inside chat
- Suggested “morning coffee” entry tasks:
- starting tasks you don’t know where to begin
- quick email editing
- screenshot analysis
- voice while on the road
Not recommended for
- Serious code
- Long documentation
- Deep, single-topic in-depth work (specialists beat it)
Pros
- Broad coverage; least likely to be “wrong for the task”
- Strong tooling (plugins/custom GPT, canvas)
Cons
- Versatility becomes a minus in 2026 due to specialization
- Not the top choice for deep specialist tasks
Numerical score(s)
- 58.6 (SVE bench)
2) Claude (Anthropic) 4.7 / 4.8 (Claude Optimized for text+code)
Positioning / philosophy
- Presented as a different brand/philosophy:
- OpenAI = generalist
- Anthropic = specialist in text and code
- Claimed market outcome: developer tools (examples given: Cursor, Windsurf) work mainly on Claude.
Performance / capabilities
- Claude 4.7: 87.6 on SVE bench (called best among all models at the time; “key detail” is Claude is a specialist)
- Up to 128,000 tokens in one go (unique among “top models” per the video)
- described as being able to write a full book in one answer without losing style/logic
- Correction behavior:
- Claude is described as the only top model likely to explicitly say “you’re wrong”
- 4.8 improves this further (described as less “overly agreeable”)
Tone / UX
- Less “praise/agree/nod” behavior than GPT/Gemini
- “Annoying for first 2 hours,” then “priceless” after users realize its value for serious work
Not recommended for
- “Everyday trifles”
- Image generation (not available / not here)
- No voice communication (voice only in English, per subtitles)
Pros
- Best-in-class for professional code and serious text
- Large context output (128k tokens)
- Direct correction rather than flattery
- Better fit for long legal docs, complex analysis
Cons
- Not good for quick casual tasks (as portrayed)
- Limited voice (English only) and absent multimodal/image capability (per subtitles)
Numerical score(s)
- 87.6 (SVE bench) for Claude 4.7
- 128,000 tokens context/output claim
3) Gemini (Google) 3.1 Pro
Performance / features
- Claims:
- Best in “pure thinking tests”: 94.1 on GPQA Diamond
- 1 million token context window (upload a novel / many PDFs / long video)
- Large context use case: answer across the entire uploaded array
- Price advantage (claimed):
- $2 per million input tokens
- $12 per million output tokens
- Compared to others:
- “three times cheaper than Claude”
- “almost one and a half times cheaper than GPT 5.5” (for comparable quality)
Best use cases
- Users in the Google ecosystem (Gmail, Docs, Drive, Calendar, YouTube) get a stronger assistant because it can use that context.
- Also recommended for:
- large documents/research
- anything tied to Google services
- video analysis
- low-cost, high-volume tasks
Not recommended
- “High level code” (Claude still better)
- For “text aesthetics,” Claude is said to be ahead
Pros
- Huge context (1M tokens)
- Strong scientific/data-thinking benchmark (GPQA)
- Very cost-effective, especially for high volume
- Strong ecosystem integration (per video)
Cons
- Not the best for high-end coding (in this comparison)
- Less favored on text aesthetics
Numerical score(s)
- 94.1 (GPQA Diamond)
- 1,000,000 tokens context window (claim)
- Pricing: $2 input / $12 output per 1M tokens
4) Grok (real-time X feed access)
Unique feature
- Direct access to the X feed in real time
- Practical result: answers about what’s hot from the last ~10 minutes, not old training data.
Best use cases
- Real-time IT community trends
- News, trends, current events, social-media sentiment analysis
- “Anything relevant to my niche” (implied current events research)
Not recommended
- Tasks where neutral tone is important (corporate documents)
- Working with children
Tone / UX
- Can be joking or sharp; may deliver politically incorrect truth
- For some, a plus; for others, a reason to avoid
Performance mentions
- Pure tests: code is near the top:
- 75% on SVE bench
- Math:
- 50.7% on Humanitest Last Exem (described as very difficult)
Pros
- Best for fresh, real-time information
- Good enough coding performance (near GPT level, below Claude)
Cons
- Not reliably neutral; potentially risky tone for sensitive use
Numerical score(s)
- 75% (SVE bench)
- 50.7% (Humanitest Last Exem)
5) DeepSeek (open-source, local run)
Key features
- Open source and “fully open,” downloadable for local execution
- Claimed benefits:
- No API fees
- No corporate data leakage to other servers
- Full control
Performance / pricing
- Cited: >80% on SVE bench
- Pricing stated:
- $74 per million input tokens
- Compared with Claude OPUS 4.7 at $15 per million
- (Video implies it’s “science fiction” expensive in API terms, but makes the case that local running changes the economics.)
Best use cases
- Large volumes of code
- Any tasks where privacy is required
- Bulk data processing
Not recommended
- Creative texts
- Emotional storytelling
- Multimodal image tasks (implied not good / not supported)
Pros
- Strong privacy/control (local)
- High code performance (especially in the “bulk code” scenario)
- Economics reframed around owning a gaming PC
Cons
- Not for creative/emotional writing and image/multimodal needs (per video)
- API pricing described as high, but local use is the workaround
Numerical score(s)
- >80% on SVE bench (DeepSeek)
- $74 per million input tokens (as stated)
Media models mentioned (not a single “product”)
The video also claims specialized leaders for image/video generation.
Images
- GPT Image 2: #1 in rankings; “242-point lead” over previous leader
- best for images with text inside (posters, covers, UI, mockups)
- Google Nanobona Pro: best portraits / reference work
- up to 14 photos of a face + Google search during generation
- Nanobanana 2 Lite (4K): “in 5 seconds” (per subtitles)
Video
- “42”: best physics
- videos up to 25 seconds
- Veo 3.1: cinematic 4K with native sound generated
- CLН 3.0: first 4K at 60 fps, plus “free plan of 66 credits per day”
- ECDEN 2.0: up to 12 input files, leader for multi-scene narrative
- Runway 4.5: directorial control of camera and effects
Overarching media verdict
- No universal approach in media; each tool covers classes of tasks and trying to do everything with one model fails.
The “$2,000/year” strategy (core recommendation)
The video’s promised pattern:
- Use one paid base account for about 80% of tasks:
- often Claude (for techies) or Gemini (for Google ecosystem)
- Add free access for targeted tasks:
- Grok via Perplexity (basic plan)
- DeepSeek via web interface (“webinfйс”)
- Cost framing:
- $12/month instead of $80 (as stated)
Important caveat: Which model is “leader” can change every 2–3 weeks, so users should keep checking updates.
Pros vs Cons (as presented overall)
Pros of the strategy
- Higher quality per task via specialization
- Significant cost savings (~$2,000/year claim)
- Better fit for real workflows (docs, code, real-time info, bulk local privacy)
Cons / risks
- Requires model switching and monitoring (leaders change frequently)
- Not every model covers every modality (e.g., Claude lacks voice/image per subtitles; Grok tone can be risky)
Verdict (concise)
Recommended approach: Build a small stack rather than paying for a single “universal” model.
- Pick Claude if your work is heavy coding + serious text.
- Pick Gemini if you’re in the Google ecosystem and need huge context + cost efficiency.
- Add Grok for real-time X/news and DeepSeek for privacy/bulk code (often locally).
Overall, the video strongly supports the “right model for the right job” plan and claims it can save ~$2,000/year.
Speakers / perspectives
- Single speaker throughout (no separate speakers identified in the subtitles). All comparisons and recommendations appear to come from the same tester/narrator describing their 3-week evaluation across models.