Video summary

Отдал ИИ переписку отдела продаж — он нашёл 1,8 млн потерь

Main summary

Key takeaways

Technology

What the video demonstrates (tech + product capability)

  • The author builds an AI-based CRM quality-control / sales auditing system using LLM agents.
  • Two “controllers” are tested to automatically:
    • Read AMCRM deal cards (and also call transcripts).
    • Detect rule violations according to a Sales Manager CRM implementation standard (a checklist with numbered rules).
    • Write results back into each deal card as comments with exact point numbers and verbatim quotes from correspondence/calls.
    • Identify violations such as:
      • missing clarification
      • wrong pricing without needs discovery
      • lack of next-step assignment
      • lack of loss reasons
      • weak handling of objections
      • sarcastic/devaluing responses

Tools / setup described

  • Uses Visual Studio Code (free) as the IDE.
  • Installs two VS Code extensions (LLM integrations):
    • OpenAI / Codex (GPT-5.6 “Soul” in the narration)
    • Anthropic Cloud / Claude (called “Claude Fable 5”)
  • Notes the need for paid subscriptions for model execution (approx. pricing given):
    • OpenAI via tokens / cloud usage
    • Claude via a subscription level
  • Uses a “Bypass Permissions” / full-access-like mode (agents decide actions without extra approvals), with an explicit safety warning that it should only be enabled when well understood.
  • Access to AMCRM is handled via an API access file (the author mentions retrieving access data from AMCRM integration settings).
  • Implements anonymization to comply with Federal Law 152 (e.g., removing/obscuring personal data like names/phones/emails before sending to the model).

Evaluation methodology (battle format)

  • Two models are compared head-to-head on the same task.
  • Tests run in multiple stages:
    1. Models ask questions and propose a plan (before writing code)
    2. Audit 1 transaction
    3. Audit 5 complex transactions
    4. Audit all 40 transactions at once
  • Scoring is based on:
    • correctness of detected violations
    • proper attention to call transcripts, not just chat messages
    • whether the model follows the written standard precisely (by point number)
    • efficiency and completeness

Key findings from the audit results

  • Both models generally work and produce numbered violations aligned to the author’s standard.
  • Early-stage:
    • Both successfully complied with anonymization rules and understood the task.
    • “Soul/Codex” sometimes missed/handled some outputs differently (e.g., notes/recording details), while “Fable/Claude” caught violations quickly.

5-transaction round (highlights)

  • Fable (Claude) was often faster.
  • Soul (Codex) often found more violations and/or produced stricter audit outputs (more “picky”).
  • Example outcomes described:
    • In some deals, Soul found additional issues in cards/correspondence.
    • In others, Fable better identified context via chats/calls or applied certain points more appropriately.

Full 40-transaction run

  • The author used a script to compare model outputs against his manual “truth” markings.
  • Reported accuracy comparison:
    • Soul: 40/40 (strong overall alignment with what the author later verified)
    • Fable: 12/40 in the strict comparison, but closer to the author’s softer “trap” expectations
  • “Trap” / no-violation setup:
    • The author claims one “clean” deal (and other specifics) should not be penalized; the goal was to detect hallucinated/overconfident “fake” violations.
    • Fable behaved more gently/realistically on these traps; Soul was harsher but still sometimes fair.

Monetization / ROI claim (loss estimation)

From the audit of completed and in-progress deals, the author states:

  • Sales department losses: ~1.815 million rubles
  • Plus 1.515 million rubles in deals currently “in progress” (still potentially salvageable)

He also compares closed-lost vs closed-won patterns:

  • Closed deals with low violations vs lost deals with high violations per transaction.
  • This is used to estimate how much money leaks from process violations.

Deliverables / “productized” output

  • Each model generates files/reports (one folder per model) containing audits per deal.
  • The author shows one model can draft a daily letter to the sales head summarizing:
    • total violations
    • where money was lost
    • per-employee / per-deal issues
    • which deals can be saved
  • He states he prefers one letter’s quality, but overall declares:
    • Result: draw (2:2)
    • Suggests both could be improved with iterations.

Practical takeaway / guide aspect

  • The author says the CRM implementation standard + 40 test transactions (and supporting files) will be posted on Telegram after publishing.
  • He argues the standard can be used without AI:
    • Manually run 5–10 deals through the checklist to find where money is leaking.
  • Encourages viewers to propose the next business area for AI “battle” testing.

Main speakers / sources

  • Primary speaker: the video author (“I” throughout), conducting the CRM audit experiment and explaining the setup.
  • AI sources tested:
    • Claude (“Claude Fable 5”)
    • OpenAI Codex / GPT-5.6 (“Soul”)

Original video