Video summary
The Ultimate Beginner’s Guide to Hermes Agent
Main summary
Key takeaways
Summary
The video is a beginner-to-advanced guide for “Hermes Agent”, positioning it as an AI “employee” / chief-of-staff that runs on a computer and can operate across tools, channels, and workflows—becoming more useful over time via self-improvement memory loops and reusable skills/workflows.
1) What Hermes Agent is (and why it’s different)
- Hermes is described as a “harness”/runtime around an AI model (e.g., “GPT 5.5”, Claude Opus, etc.).
- The core differentiator vs other agents/harnesses (examples mentioned: OpenClaw, Cloud Code, Codeex) is:
- Hermes behaves like an AI employee that grows with you, not a one-off tool.
- It uses a self-improvement loop where repeated successful work is stored as a reusable skill, so future tasks execute closer to your preferences automatically.
- The video emphasizes Hermes becomes more reliable over time by:
- improving/archiving skills strategically,
- avoiding “bloated” or breaking agent behavior.
2) Fundamentals before setup: security, cost, and limitations
- An analogy is used:
- Model = engine
- Hermes = car harness enabling computer control, channels, and memory
- Claims:
- Model-agnostic design
- Local data: information stays on the user’s machine to reduce vendor lock-in
- Key caution:
- These agents are not fully autonomous decision makers
- Users must:
- limit permissions
- verify outputs
- add safeguards before high-risk actions (email, money movement, publishing)
- Example risk mentioned:
- an agent with email access allegedly deleted years of emails
- takeaway: controlled access + verification is essential
3) Install & set up Hermes (two deployment paths)
- Two deployment options:
- VPS/cloud server
- Local machine (e.g., Mac Mini / old PC)
- The demo focuses on Mac setup.
- “Buddy system” concept:
- Use Codeex as a developer buddy/doctor to troubleshoot Hermes issues.
- Suggested as “insurance” when setup breaks or remote help is needed.
- Claim: Codeex can securely interact with apps from a phone even if the Mac is locked/off.
4) Model configuration and backups
Hermes desktop setup includes selecting:
- Primary model
- Recommended example: OpenAI / Codeex-branded plan using GPT 5.5 via subscription
- Fallback/backup model via Newsportal
- Described as a “backup brain” with access to top-tier models and audio models
- Mentions:
- “fast mode”
- cost optimization via a cost-effective model mix
5) “Computer Use” integration (controlling your desktop)
Hermes “computer use” is presented as a capability for the agent to:
- install required drivers/packages,
- request permissions (including “KUA driver access” in subtitles),
- then control the computer after access is granted—enabling hands-free task execution.
6) Multi-channel “gateways” (Telegram, Slack, iMessage, etc.)
The tutorial stresses connecting messaging channels first because they enable:
- measuring work done while away,
- tracking automation percentage,
- treating Hermes like an “AI employee” and measuring “revenue per employee”.
Telegram
Hermes supports patterns such as:
- a direct solo thread (Hermes replies),
- a capture/ingest channel where Hermes can remember items without needing user responses.
Topic chats are demonstrated using:
- BotFather,
- bot tokens,
- pairing codes,
- admin permissions.
Topics enable multiple parallel threads (tabs) and show when Hermes is actively typing.
Slack
- Configure using:
- bot token
- Slack app creation
- Mentions:
- socket mode
- a “gateway restart” concept to ensure Hermes can talk to the channel
- Supports team context and optionally whitelisting users.
iMessage
- Uses Photon.codes to bridge iMessage into the Hermes gateway.
- Phone messaging to Hermes is enabled after creating/applying credentials for a configured project number.
7) Identity + personality (“soul”, profiles, identity files)
Hermes personality is driven by identity files:
- A “soul” file controls tone, style, and purpose
- example behaviors: lower-cased, concise, “in-line”
- An interview flow can generate user identity data (e.g., a “user.md” concept)
“Profiles” let users create multiple Hermes personalities/agents with different configurations:
- roles,
- models,
- memory,
- skills,
- tool access.
8) Advanced memory: why multi-session agent memory is hard
The video frames memory as a complex systems problem:
- Single-session memory
- maintaining preferences and commitments within one long conversation
- Knowledge updates
- not treating old facts as still-current truth
- Multi-session memory
- connecting context across Telegram/Slack/other sessions
- Temporal reasoning
- determining what changed over time and what is now outdated
9) Honcho as a “shared brain” / memory router
Honcho is introduced as a memory layer that:
- addresses “cold start” by routing/maintaining context across chats,
- creates peer cards (a diacronic identity concept) representing relationships with different people/agents,
- uses reasoning models (mentioned: “neuromancer”) to derive contradictions/implications.
Safety note:
- Honcho is described as dangerous if given too much power,
- but can be self-hosted for privacy.
Integration:
- Hermes can use Honcho as its memory provider (via OAuth).
10) Agent management: Multica mission control + hooks
Multica is presented as an observability + orchestration layer:
- manage multiple “runtimes” and workspaces,
- track issues/projects/tasks handled by Hermes and other AI employees.
Hooks are used to create deterministic triggers, such as:
- automatically updating Multica projects/issues when Hermes finishes tasks.
Demo theme:
- Hermes acts like an AI project manager / chief of staff moving work across a board without manual saving.
11) Building blocks for real integrations: MCP, CLI, API, and connectors
Connectors communicate with external apps via interfaces like:
- MCP
- CLI
- API
Examples:
- Composio (MCP) as a gateway for many apps (Google Calendar/Gmail, Slack, HubSpot, Twitter/X, etc.)
- Google Workspace CLI for Drive/Docs/Slides editing
- cron jobs for recurring tasks
12) Cron jobs & “morning briefing” autopilot
A recurring automation is demonstrated that:
- pulls calendar meetings,
- summarizes urgent emails,
- reads Slack team context,
- finds AI trending topics (e.g., Twitter),
- delivers a short practical briefing to Telegram.
Concepts used:
- skills: turn repeated workflows (morning briefing format) into reusable units
- agentic workflows: loops like research → curate → draft/review → deliver (not single prompts)
13) Build #2: automated animated website generation (“/slashgoal”)
Introduces /goal (slashgoal):
- sets an outcome goal,
- drives an iterative loop across turns until the goal is achieved.
Supporting tools/skills mentioned:
- Higsfield AI (image/photo generation + deploy)
- Firecrawl (scrape competitor sites)
- here.now (deploy/hosting)
Outcome:
- Hermes builds a premium animated landing page (graphics, copy, media elements, self-checking/refinement)
- delivered as a deployed URL
14) Build #3: monetizing with “property marketing kits” (realtors)
Offer workflow:
- scrape properties (Firecrawl),
- generate marketing kits (upgraded visuals + room reconstructions),
- compile sources and deliver assets/contact info for outreach.
Mentions:
- a prebuilt skill in “cloud club”:
- property marketing kit (described as a cloud-code skill)
Pricing (approximate ranges mentioned in subtitles):
- about $250–$1,000 per listing
15) Build #4: content agent / growth automation
A content pipeline spans:
- research trends/keywords,
- scriptwriting,
- filming/producing,
- editing,
- posting,
- analyzing.
Emphasis includes:
- carousels (Twitter/X-style and styled variants),
- prebuilt skills for “Instagram thread carousel” styles,
- a “content engine” automation driven by cron jobs and posting integrations.
Posting automation includes:
- Composio
- Blotato for scheduled posting calendars (noted: OAuth configuration needed)
“Content octopus” concept:
- repurpose one YouTube video into shorts, articles, tweets, LinkedIn lead magnets/posts, carousels, etc.
- uses a content repurpose skill
16) Video editing automation with skills
Demonstrates editing a raw talking-head clip via a two-phase approach:
- Cut video skill
- removes filler/poor takes
- B-roll finder skill
- finds B-roll and inserts it
Result:
- a polished reel-like output.
17) Ecommerce ad factory + competitor ad inspiration
Ad creative generation approach:
- use competitor research to guide messaging,
- generate images/videos,
- insert the creator into ads using a generated identity.
Mentions:
- deploying/optimizing via an ads platform integration (example: Meta) at scale.
18) Community/customer support agent + inbox automation
A “school manager” / community manager agent:
- checks new posts/comments via cron,
- summarizes sentiment,
- answers questions.
Integration:
- Agentmail creates an agent mailbox:
- create inbox,
- provide Hermes credentials,
- send/receive support emails,
- manage support tickets.
19) Build: trading bot with self-improvement behavior
Trading bot framed as self-improving via Hermes self-improvement loops and skills.
Architecture:
- broker/trading platform: co-invest
- supports paper trading and switching to real trading
- strategy/data source: QuiverQuant
- congressional trading + strategies
- connectors via MCP
Workflow demonstrated:
- connect QuiverQuant MCP,
- copy trade a person/strategy,
- use cron jobs during trading hours,
- generate weekly dashboards via here.now
20) Build: longevity/health coach with daily check-ins
Health workflow:
- ingest blood panel PDF,
- ingest sleep/workout sources (e.g., Oura/AW8 mentioned in subtitles),
- agent produces a supplement plan,
- daily reminders via cron,
- voice memos to reduce friction and increase accountability.
21) “Jarvis mode” and a mission-control UI
A “Jarvis” command center includes:
- calendar organization,
- task agenda (“top three things”),
- computer control,
- voice output via 11 Labs (or local alternative, “voice box” mentioned).
Built by combining:
- honcho for life context memory,
- Composeio + Google MCP for calendar automation,
- computer use,
- voice,
- HUD/UI built with design/coding orchestration (references include Claude + Fable + Codeex).
Theme: chief-of-staff UI + voice interaction.
22) Business + selling: pricing framework (productized services)
Go-to-market summary:
- “Price the outcome, not hours.”
- Recommended offer ladder (example pricing mentioned):
- AI employee install (non-technical): ~$1,500 setup (+ optional retainer)
- Executive assistant / autopilot (cron + automation): ~$500
- Animated website lead magnet/service: ~$250–$500
- Property marketing kits: ~$250 per listing (optionally scaled via autopilot)
- Content engine (carousels + posting): positioned as high leverage
- Content repurpose (weekly long-form → multi-platform): ~$1,000–$3,000/month
- Ad factory: creatives generation + scaling via ad platform integrations
- Community manager/customer support: ~$500 or per incident
- Trading bot + longevity coach: positioned as higher-ticket premium work
- Jarvis/Agent Club demo/performative product for closing calls
Advice emphasized:
- demo and build first for yourself,
- use a risk-reversal approach (e.g., ensure the client doesn’t bear most setup risk).
Main speakers/sources
- Main speaker: the course creator/instructor (first-person throughout; name not provided in subtitles)
- Referenced external source: Mark Cuban (motivational/business framing clip included)