Video summary
Claude Code Harness 심화 / 황민호 / 260527
Main summary
Key takeaways
Overview of the session
- The video is a deep-dive training on “Harness Engineering” (hänes/harness; auto-captioned inconsistently), taught by Hwang Min-ho (Kakao) as an extension of earlier Claude Code basics/training.
- The speaker frames harness engineering as moving from manual “steering” with AI toward letting AI run longer autonomous workflows by building an agent team + orchestration + verification loop.
Main technological concepts explained
1) What “Harness Engineering” is (definition + intuition)
- Harness is described as a system that configures AI agents so they can perform many tasks over long periods with minimal human intervention.
- The term is likened to a horse harness: AI’s performance depends on how you structure/configure its harness.
- Key emphasis: harness engineering is not just better model choice, but how the model is wrapped with:
- environment
- workflow
- tools
- agents
- checks
2) Multi-agent design + orchestration workflow
The harness is built from key components:
- Agents
- defined with role descriptions and protocols in natural language
- Skills
- reusable procedures/checklists/tools with structured inputs/outputs
- Orchestrator (PM-like coordinator)
- schedules steps
- monitors ordering/progress
- prevents agents from going off-track
- triggers rework when failures occur
Core execution loop:
- Decompose task
- Distribute to agents (often in parallel) → produce deliverables
- Test/QA → feedback → iterate until quality targets are met
Kitchen analogy (as presented):
- Single agent: works until a bottleneck forms
- Multi-agent: parallel work + maintained quality
- Tradeoff: multi-agent is more scalable but costs more tokens/time
3) Human responsibility remains essential
Even with harnesses, the speaker stresses:
- Humans must define goals
- Humans must manage context
- Humans must specify/enforce quality criteria
- Humans remain responsible for ethics/responsibility
- Outputs are treated as initial results that still require refinement and review
4) Context management (“workbench” idea)
- Context is compared to ingredients on a kitchen workbench:
- irrelevant/mixed context increases errors
- conflicting internal thoughts confuse the system
- The orchestrator is presented as the mechanism that ensures tasks happen in the correct order with correct context.
Product/features and tooling demonstrated (Claude Code + plugins + harness skills)
1) Claude Code training prerequisites + plan/token guidance
- The session mentions:
- potential token limits for Pro Annual plan users
- recommendation to switch to Max plan if tokens run out during training/exercises
- A recap of earlier Claude Code work included:
- terminal basics
- creating web/skills/agents
- context verification and token/context management topics
2) Installing a “Meta Harness” skill (hands-on tutorial)
A practical step-by-step flow is shown:
- Add a Marketplace (auto-captioned names include “Slush Plugins / Rev Factory / Slush Harness”, etc.)
- Install the Harness plugin
- If needed, install a reload plugin so “Slash Harness” becomes runnable
Failure mode mentioned:
- Sometimes the plugin may not install cleanly; use the “slash plugin” method to install via the Marketplace UI.
3) Using Slash Harness / autocomplete “harness”
Workflow UI behavior:
- Typing “harness” + Tab triggers autocomplete for harness configuration
/Slash Harnessregenerates/updates harness configuration with a new intention/goal
Demonstrated pattern:
- agents are created that perform:
- research → analysis → report writing → QA review
- based on stated intent (e.g., “my intention is to investigate/write…”), rather than “write a report directly”
4) Agent visibility / Auto Mode / Plan Mode / Agent Team Mode
Mode switching (via Shift+Tab) is discussed:
- Auto Mode
- helps the system proceed without repeated human confirmations
- Plan vs Access-like modes are also mentioned (auto-captioned)
Known issue:
- The agent team sometimes fails to load dynamically, showing regular agents instead of team agents
- Fix: reloading/restarting resolves it
- Team-mode works more reliably when loaded dynamically
Later Q&A clarifications:
- Intermediate deliverables are checked for structure/quality
- Token efficiency is managed by review and controlling intermediate outputs (not fully automated by hiding all detail)
5) Cross-tool verification: Claude + Codex CLI (reducing bias/hallucination)
QA is upgraded conceptually:
- Primary verification with Claude
- Secondary verification with Codex CLI
An updated orchestrator/QA agent configuration is shown where:
- QA uses Codex CLI for verification tasks
Motivation given:
- single verification can be biased; combining verifications improves trustworthiness
Concrete use cases shown as examples
1) Government policy / R&D budget analysis (demo project)
A large demo harness includes:
- agents for budget allocation analysis
- agents for strategy under PBS / post-PBS system reform
- agents for policy research across document sources (press releases, opinions, agenda items, etc.)
- a report-writing agent and a QA reviewer
The demo shows intermediate workspace folders holding outputs such as:
- policy overviews
- sector breakdowns
- “problems/limitations” with traceable sources
- QA reports with revision suggestions
Final output demo includes:
- scenario-based analysis (downward / neutral / conservative)
- visualization via a generated webpage + simulation elements
- a note that PDF download features had some issues during the demo
2) Webtoon generation workflow (multi-agent creative pipeline)
Harness is shown generating:
- novel outlines/plans from a logline
- character design agent
- world-building agent
- science verification agent that revises text to match algorithm-readable constraints
Expansion to webtoons:
- “webtoon artist” agent draws pages
- “quality verification agent” checks:
- character consistency
- image correctness
- speech bubble rendering
- If QA fails, it triggers redraw (loop)
- Iteration with user feedback improves dialogue/image connections
3) Markdown viewer + privacy/storage approach
A demo “Markdown viewer” is built that:
- supports HWP import (not available in other markdown viewers, per the speaker)
- stores documents only in the browser, not transmitted to servers (privacy consideration)
4) Data center research + media sentiment + hallucination reduction
Harness is used to analyze a newly opened national AI data center:
- organizes a timeline
- analyzes both critical and positive media outlets
- evaluates topics such as:
- business viability
- GPU supply/demand
- power environment
- data sovereignty
Claim:
- using source-driven, structured citations reduces hallucinations significantly (though not to zero).
Review / performance evaluation claim
- The speaker describes an experiment:
- created 15 development tasks (including a key-value store example)
- compared performance with vs without harness
- used an LLM judge for evaluation
- Reported results:
- ~60% average performance improvement
- when forming a “team for a year” (captioned oddly), win rate reported as 100%
- for harder tasks, simple AI tool use drops, while harness maintains quality
Limitations, tradeoffs, and operational concerns mentioned
- Multi-agent harness:
- increases token usage, especially during agent-to-agent communication (speaker cites ~30% increase in measurements)
- runs slower under heavy use; Claude Code speed concerns are discussed (speed benchmarks referenced)
- Tool competition/updates:
- speaker compares speed/quality between Claude Code and Gemini variants in a live website-building comparison
- Gemini is described as faster and able to generate images (with a note that Claude can also generate images via a skill)
Main speakers / sources (as stated)
- Engineer Hwang Min-ho (Kakao) — primary speaker/author of the training.
- “Kim Bok-eum from the Strategy Office” — session host who introduces the speaker and logistics (IP/LAN adapter, token plan advice, feedback collection).