Video summary

Claude Code Harness 심화 / 황민호 / 260527

Main summary

Key takeaways

Technology

Overview of the session

  • The video is a deep-dive training on “Harness Engineering” (hänes/harness; auto-captioned inconsistently), taught by Hwang Min-ho (Kakao) as an extension of earlier Claude Code basics/training.
  • The speaker frames harness engineering as moving from manual “steering” with AI toward letting AI run longer autonomous workflows by building an agent team + orchestration + verification loop.

Main technological concepts explained

1) What “Harness Engineering” is (definition + intuition)

  • Harness is described as a system that configures AI agents so they can perform many tasks over long periods with minimal human intervention.
  • The term is likened to a horse harness: AI’s performance depends on how you structure/configure its harness.
  • Key emphasis: harness engineering is not just better model choice, but how the model is wrapped with:
    • environment
    • workflow
    • tools
    • agents
    • checks

2) Multi-agent design + orchestration workflow

The harness is built from key components:

  • Agents
    • defined with role descriptions and protocols in natural language
  • Skills
    • reusable procedures/checklists/tools with structured inputs/outputs
  • Orchestrator (PM-like coordinator)
    • schedules steps
    • monitors ordering/progress
    • prevents agents from going off-track
    • triggers rework when failures occur

Core execution loop:

  1. Decompose task
  2. Distribute to agents (often in parallel) → produce deliverables
  3. Test/QA → feedback → iterate until quality targets are met

Kitchen analogy (as presented):

  • Single agent: works until a bottleneck forms
  • Multi-agent: parallel work + maintained quality
  • Tradeoff: multi-agent is more scalable but costs more tokens/time

3) Human responsibility remains essential

Even with harnesses, the speaker stresses:

  • Humans must define goals
  • Humans must manage context
  • Humans must specify/enforce quality criteria
  • Humans remain responsible for ethics/responsibility
  • Outputs are treated as initial results that still require refinement and review

4) Context management (“workbench” idea)

  • Context is compared to ingredients on a kitchen workbench:
    • irrelevant/mixed context increases errors
    • conflicting internal thoughts confuse the system
  • The orchestrator is presented as the mechanism that ensures tasks happen in the correct order with correct context.

Product/features and tooling demonstrated (Claude Code + plugins + harness skills)

1) Claude Code training prerequisites + plan/token guidance

  • The session mentions:
    • potential token limits for Pro Annual plan users
    • recommendation to switch to Max plan if tokens run out during training/exercises
  • A recap of earlier Claude Code work included:
    • terminal basics
    • creating web/skills/agents
    • context verification and token/context management topics

2) Installing a “Meta Harness” skill (hands-on tutorial)

A practical step-by-step flow is shown:

  • Add a Marketplace (auto-captioned names include “Slush Plugins / Rev Factory / Slush Harness”, etc.)
  • Install the Harness plugin
  • If needed, install a reload plugin so “Slash Harness” becomes runnable

Failure mode mentioned:

  • Sometimes the plugin may not install cleanly; use the “slash plugin” method to install via the Marketplace UI.

3) Using Slash Harness / autocomplete “harness”

Workflow UI behavior:

  • Typing “harness” + Tab triggers autocomplete for harness configuration
  • /Slash Harness regenerates/updates harness configuration with a new intention/goal

Demonstrated pattern:

  • agents are created that perform:
    • research → analysis → report writing → QA review
  • based on stated intent (e.g., “my intention is to investigate/write…”), rather than “write a report directly”

4) Agent visibility / Auto Mode / Plan Mode / Agent Team Mode

Mode switching (via Shift+Tab) is discussed:

  • Auto Mode
    • helps the system proceed without repeated human confirmations
  • Plan vs Access-like modes are also mentioned (auto-captioned)

Known issue:

  • The agent team sometimes fails to load dynamically, showing regular agents instead of team agents
  • Fix: reloading/restarting resolves it
  • Team-mode works more reliably when loaded dynamically

Later Q&A clarifications:

  • Intermediate deliverables are checked for structure/quality
  • Token efficiency is managed by review and controlling intermediate outputs (not fully automated by hiding all detail)

5) Cross-tool verification: Claude + Codex CLI (reducing bias/hallucination)

QA is upgraded conceptually:

  • Primary verification with Claude
  • Secondary verification with Codex CLI

An updated orchestrator/QA agent configuration is shown where:

  • QA uses Codex CLI for verification tasks

Motivation given:

  • single verification can be biased; combining verifications improves trustworthiness

Concrete use cases shown as examples

1) Government policy / R&D budget analysis (demo project)

A large demo harness includes:

  • agents for budget allocation analysis
  • agents for strategy under PBS / post-PBS system reform
  • agents for policy research across document sources (press releases, opinions, agenda items, etc.)
  • a report-writing agent and a QA reviewer

The demo shows intermediate workspace folders holding outputs such as:

  • policy overviews
  • sector breakdowns
  • “problems/limitations” with traceable sources
  • QA reports with revision suggestions

Final output demo includes:

  • scenario-based analysis (downward / neutral / conservative)
  • visualization via a generated webpage + simulation elements
  • a note that PDF download features had some issues during the demo

2) Webtoon generation workflow (multi-agent creative pipeline)

Harness is shown generating:

  • novel outlines/plans from a logline
  • character design agent
  • world-building agent
  • science verification agent that revises text to match algorithm-readable constraints

Expansion to webtoons:

  • “webtoon artist” agent draws pages
  • “quality verification agent” checks:
    • character consistency
    • image correctness
    • speech bubble rendering
  • If QA fails, it triggers redraw (loop)
  • Iteration with user feedback improves dialogue/image connections

3) Markdown viewer + privacy/storage approach

A demo “Markdown viewer” is built that:

  • supports HWP import (not available in other markdown viewers, per the speaker)
  • stores documents only in the browser, not transmitted to servers (privacy consideration)

4) Data center research + media sentiment + hallucination reduction

Harness is used to analyze a newly opened national AI data center:

  • organizes a timeline
  • analyzes both critical and positive media outlets
  • evaluates topics such as:
    • business viability
    • GPU supply/demand
    • power environment
    • data sovereignty

Claim:

  • using source-driven, structured citations reduces hallucinations significantly (though not to zero).

Review / performance evaluation claim

  • The speaker describes an experiment:
    • created 15 development tasks (including a key-value store example)
    • compared performance with vs without harness
    • used an LLM judge for evaluation
  • Reported results:
    • ~60% average performance improvement
    • when forming a “team for a year” (captioned oddly), win rate reported as 100%
    • for harder tasks, simple AI tool use drops, while harness maintains quality

Limitations, tradeoffs, and operational concerns mentioned

  • Multi-agent harness:
    • increases token usage, especially during agent-to-agent communication (speaker cites ~30% increase in measurements)
    • runs slower under heavy use; Claude Code speed concerns are discussed (speed benchmarks referenced)
  • Tool competition/updates:
    • speaker compares speed/quality between Claude Code and Gemini variants in a live website-building comparison
    • Gemini is described as faster and able to generate images (with a note that Claude can also generate images via a skill)

Main speakers / sources (as stated)

  • Engineer Hwang Min-ho (Kakao) — primary speaker/author of the training.
  • “Kim Bok-eum from the Strategy Office” — session host who introduces the speaker and logistics (IP/LAN adapter, token plan advice, feedback collection).

Original video