Video summary

You've Been Paying For The Wrong Half Of AI

Main summary

Key takeaways

Technology

Main technological concept

  • DeepSeek Harness: a “harness” (agent framework) that wraps a selected model (“the brain”) with additional capabilities such as:

    • tools
    • MCP servers
    • permissions
    • memory
  • Key distinction vs Claude Code: Claude Code is positioned as a harness where many internal decisions are closed/fixed by someone else. DeepSeek Harness, by contrast, keeps those choices open at runtime, so the agent can adapt without relying on hardcoded assumptions.

Why it’s claimed to be different/better

  • Component dependency resilience

    • If you remove a component or a dependency fails mid-task, the agent can break instead of restarting—unless a contingency plan was hardcoded.
    • DeepSeek Harness aims to avoid hardcoding by having components declare what they need, with runtime support to adjust behavior (described as “undo button” behavior).
  • Formal foundation mentioned

    • A referenced paper and a framework called Cordis are cited as the underlying idea, though no details are provided in the summary.

Installation / setup tutorial (explicit steps)

  • The video provides a step-by-step install and run guide.

Two installation options

  1. Website (described as “everything is a plugin”)
  2. GitHub repo (noted as extremely popular, with 200,000+ stars)

Local run flow

  • The harness installs successfully after waiting.
  • Open the localhost dashboard.
  • It initially wasn’t running, so the speaker starts it from the dashboard.
  • The dashboard initially prompts for an API key, which can be configured later.

Dashboard features highlighted

  • Settings → Models → Add provider

    • Example shown: add the Anthropic provider and paste an Anthropic API key.
    • The model provider becomes active after applying changes.
  • “Everything is a plugin” architecture

    • The dashboard UI is itself plugin-based (commands, sidebar, etc.).
    • Supports adding, enabling, disabling, and deleting plugins.
  • Creator mode

    • Enables plugin development skill, allowing users to build custom plugins for the harness.

“Run it for free” guidance + limits

  • The speaker emphasizes:
    • The harness is free “forever”.
    • The potentially paid part is the model API, since cost comes from the “brain” provider.

Truly uncapped free usage

  • Run the model locally
  • Mentions Ollama as an easy one-line install, but notes the speaker can’t demo it on their specific hardware in the video.

Caveats

  • Some hosted models may work via free tiers (example: Gemini), but limits are not guaranteed and providers may stop publishing limits.
  • The “genuine uncapped free path” is local inference.

Demo: building and using a custom plugin

  • The speaker creates a plugin for a practical quality-of-life feature:
    • Desktop notification when a run finishes
    • Notification includes token count and cost

Demonstrated behavior

  • The run completes successfully.
  • The plugin operates correctly (working plugin in the live environment).
  • Plugins can be managed directly from the UI.

Observability / tracing and resumability

  • The video highlights full trajectory/traceability:

    • View every step categorized by type.
    • Each call is traceable, including:
      • system prompt
      • context used
      • assistant answers
      • tool invocations
      • everything between calls
  • Mid-session behavior

    • An agent is queued mid-session, then stopped, then resumed.
    • Resume preserves the same context/prompts because the harness maintains state/context.

Comparison vs Claude Code / Codex tools (important details)

  • The repo reportedly includes two less-discussed packages:
    • dash subagent Claude
    • dash subagent Codex

How it works

  • The agent can call these tools to spawn child processes running Claude Code or Codex.
  • They behave like additional tools (e.g., subagent Claude, subagent Codex) that the agent can delegate to.

Limitations called out

  1. One-shot: each call starts a fresh process/conversation; it can’t be resumed.
  2. Providers dormant initially: not automatically enabled; presets decide whether the tool is used.
  3. Credential stripping: environment variables/credentials resembling keys are removed so the child process doesn’t inherit shell credentials; keys must be passed explicitly.

Framing by the speaker

  • The speaker frames this as not a direct competitor to paid agents.
  • Instead, it’s a routing layer that can drive other tools/models the user already pays for.

Reframing the “versus” narrative

  • The speaker argues the real comparison is not “DeepSeek vs Claude Code” because:

    • the model and the harness are separate purchases
    • DeepSeek Harness reduces the harness portion of the stack toward zero cost
  • Providers mentioned as shipping with DeepSeek Harness include:

    • Anthropic, OpenAI, Bedrock, Vertex, Codex
  • Example claim:
    • you can run Opus 5 inside the free harness with trace + plugins (subject to model API availability/cost).

Product status / caution

  • Labeled as developer preview / beta.
  • The README is said to warn about breaking changes and continued evolution over time.

Main speakers / sources

  • Speaker: the unnamed narrator/creator of the YouTube video (“in this video, I’ll show you…”).
  • Source referenced: DeepSeek Harness (official website + GitHub repository) and its associated documentation/README.

Original video