Video summary
Отдал ИИ переписку отдела продаж — он нашёл 1,8 млн потерь
Main summary
Key takeaways
What the video demonstrates (tech + product capability)
- The author builds an AI-based CRM quality-control / sales auditing system using LLM agents.
- Two “controllers” are tested to automatically:
- Read AMCRM deal cards (and also call transcripts).
- Detect rule violations according to a Sales Manager CRM implementation standard (a checklist with numbered rules).
- Write results back into each deal card as comments with exact point numbers and verbatim quotes from correspondence/calls.
- Identify violations such as:
- missing clarification
- wrong pricing without needs discovery
- lack of next-step assignment
- lack of loss reasons
- weak handling of objections
- sarcastic/devaluing responses
Tools / setup described
- Uses Visual Studio Code (free) as the IDE.
- Installs two VS Code extensions (LLM integrations):
- OpenAI / Codex (GPT-5.6 “Soul” in the narration)
- Anthropic Cloud / Claude (called “Claude Fable 5”)
- Notes the need for paid subscriptions for model execution (approx. pricing given):
- OpenAI via tokens / cloud usage
- Claude via a subscription level
- Uses a “Bypass Permissions” / full-access-like mode (agents decide actions without extra approvals), with an explicit safety warning that it should only be enabled when well understood.
- Access to AMCRM is handled via an API access file (the author mentions retrieving access data from AMCRM integration settings).
- Implements anonymization to comply with Federal Law 152 (e.g., removing/obscuring personal data like names/phones/emails before sending to the model).
Evaluation methodology (battle format)
- Two models are compared head-to-head on the same task.
- Tests run in multiple stages:
- Models ask questions and propose a plan (before writing code)
- Audit 1 transaction
- Audit 5 complex transactions
- Audit all 40 transactions at once
- Scoring is based on:
- correctness of detected violations
- proper attention to call transcripts, not just chat messages
- whether the model follows the written standard precisely (by point number)
- efficiency and completeness
Key findings from the audit results
- Both models generally work and produce numbered violations aligned to the author’s standard.
- Early-stage:
- Both successfully complied with anonymization rules and understood the task.
- “Soul/Codex” sometimes missed/handled some outputs differently (e.g., notes/recording details), while “Fable/Claude” caught violations quickly.
5-transaction round (highlights)
- Fable (Claude) was often faster.
- Soul (Codex) often found more violations and/or produced stricter audit outputs (more “picky”).
- Example outcomes described:
- In some deals, Soul found additional issues in cards/correspondence.
- In others, Fable better identified context via chats/calls or applied certain points more appropriately.
Full 40-transaction run
- The author used a script to compare model outputs against his manual “truth” markings.
- Reported accuracy comparison:
- Soul: 40/40 (strong overall alignment with what the author later verified)
- Fable: 12/40 in the strict comparison, but closer to the author’s softer “trap” expectations
- “Trap” / no-violation setup:
- The author claims one “clean” deal (and other specifics) should not be penalized; the goal was to detect hallucinated/overconfident “fake” violations.
- Fable behaved more gently/realistically on these traps; Soul was harsher but still sometimes fair.
Monetization / ROI claim (loss estimation)
From the audit of completed and in-progress deals, the author states:
- Sales department losses: ~1.815 million rubles
- Plus 1.515 million rubles in deals currently “in progress” (still potentially salvageable)
He also compares closed-lost vs closed-won patterns:
- Closed deals with low violations vs lost deals with high violations per transaction.
- This is used to estimate how much money leaks from process violations.
Deliverables / “productized” output
- Each model generates files/reports (one folder per model) containing audits per deal.
- The author shows one model can draft a daily letter to the sales head summarizing:
- total violations
- where money was lost
- per-employee / per-deal issues
- which deals can be saved
- He states he prefers one letter’s quality, but overall declares:
- Result: draw (2:2)
- Suggests both could be improved with iterations.
Practical takeaway / guide aspect
- The author says the CRM implementation standard + 40 test transactions (and supporting files) will be posted on Telegram after publishing.
- He argues the standard can be used without AI:
- Manually run 5–10 deals through the checklist to find where money is leaking.
- Encourages viewers to propose the next business area for AI “battle” testing.
Main speakers / sources
- Primary speaker: the video author (“I” throughout), conducting the CRM audit experiment and explaining the setup.
- AI sources tested:
- Claude (“Claude Fable 5”)
- OpenAI Codex / GPT-5.6 (“Soul”)