Video summary

Твой AI ОТУПЕЛ. Тебе не показалось…

Main summary

Key takeaways

News and Commentary

Overview

The video argues that claims of “AI models getting dumber” after updates are not simply user misunderstanding or normal variability. Instead, it presents a recurring, industry-wide pattern: companies may quietly change how models are delivered—through infrastructure, routing, limits, and response settings—without changing the actual model weights.

Core Claims and Reported Examples

1) The “gets worse right after release/update” cycle

The creator says a common pattern occurs each time a new flagship model launches with strong benchmarks:

  • Shortly after release, users report degradation, such as:
    • more repetition
    • more forgetfulness
    • shorter or less capable answers
    • slower responses
    • buggier behavior
  • The video claims companies then issue similar explanations and delays before acknowledging problems.

2) Perceived deterioration across years and models

The speaker cites anecdotal and community observations across multiple providers (including Anthropic and OpenAI, and models described as Claude / GPT, etc.). Reported themes include:

  • models responding worse over time
  • models being “lazier
  • limits being exhausted faster

3) Anthropic incident (Claude-related)

The video describes an Anthropic/Claude-related incident with several elements:

  • Performance drop around Aug 12
  • A 7-hour incident around July 31 affecting multiple products
  • The creator claims the degradation involved server/config issues, such as:
    • “wrong servers”
    • defective production
    • a portion of requests failing or producing wrong output
  • The product was allegedly effectively broken for weeks
  • The video says Anthropic later published an official analysis (dated around Sep 17), while stating it was not intentional quality reduction due to load.

4) Claude Code “reasoning level” and UI behavior changes

The video alleges that corporate changes affected:

  • default “reasoning level” (e.g., high → medium) to reduce long delays that made the UI seem frozen
  • additional changes purportedly causing forgetfulness and repetition through session-handling bugs

The creator claims users could observe quality drops, with acknowledgment allegedly coming only after weeks of complaints.

5) OpenAI “router” / model-swapping theory (central allegation)

A key argument is that users may be routed to cheaper or less capable models even when they pay for a flagship tier.

  • The creator frames this as a “broken router” or system bug (or a plausible cause).
  • The purpose would be consistent with why “flagship” users might still see worse behavior.

6) Broader pattern: cost control without changing the model itself

The video argues companies can reduce compute costs by tuning “knobs” in the delivery stack, without openly admitting that the user experience changed, for example:

  • routers
  • limits
  • wrappers (e.g., agent modes like Claude Code / Codex-style tools)
  • caching
  • response verbosity caps

7) Escalation example: Google “backs down”

As a “when pushed” example, the video mentions that when constraints are tightened (e.g., heavy requests burn through limits faster), users revolt and the company later adjusts or reverses course.

8) Grok 4.5 / context reduction claim (announced change)

The video also alleges that a rollout shipped with reduced context, justified via price/speed. It suggests this later, more transparent framing reframes earlier “quiet denials.”

Video Conclusion: What Viewers Should Do

The creator advises that since degradation and delivery changes appear predictable, users should reduce their exposure by:

  1. Keeping multiple models available (switching reduces dependency on one provider).
  2. Monitoring provider status pages (to quickly determine whether failures are systemic).
  3. Waiting after major releases before concluding the new behavior is stable.
  4. Treating AI quality like variable weather: pause/switch rather than assuming the user is wrong.
  5. Not assuming subscriptions guarantee capability, and instead prioritizing user skill:
    • detecting hallucinations/lying by cross-checking
    • requesting verification and sources
    • using critic agents

Underlying Thesis

Repeatedly, the video asserts that what users describe as “the model is getting dumb” is often a delivery-layer downgrade—such as routing, limits, reasoning settings, verbosity caps, or server issues—implemented to manage costs or mitigate performance issues, with acknowledgement arriving only after sustained user pressure.

Presenters or Contributors

Main presenter / creator

  • Not explicitly named in the subtitles (referred to as “bro” / the speaker).

Mentioned individuals or organizations

  • Sam Altman (OpenAI)
  • Anthropic (including mention of an official analysis)
  • Google (mentioned in the context of limits/compute behavior)
  • Reddit community / users (described as testing and complaining)
  • “Andrew” (appears in a comedic/skit segment: “what is this, Andrew? I am an accredited skinner…”)

Original video