Video summary

FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft

Main summary

Key takeaways

Technology

Summary (FinOps for AI Agents: “Who Spent All the Tokens?”)

The talk explores why AI agent token spending is often hard to trace and control—due to runaway loops, exploding context, and unbounded tool/model calls—and proposes a shift from token maxing to value maxing through token cost governance.


Core problem and why it’s hard

  • Token cost is the unit of cost (billing is tied to tokens).
  • Cost is generated at the LLM/model call boundary, but existing “token ops” tooling typically operates at the request level (e.g., model gateways with routing/downgrades, hard caps, halting).
  • In agent systems, what’s missing is the ability to control at the agent run / tool / sub-agent / context growth boundaries—i.e., the call path between agents and tools.

First-principles approach to cost control

  1. Treat tokens as both the cost unit and the value metric to optimize.
  2. Track where cost is created (at the model call boundary) and enable attribution.
  3. Use attribution to enforce policies that stop the root causes:
    • mitigate loops
    • prevent excessive context growth
    • reduce irrelevant tool outputs
  4. Use hard budget halting only as a last resort (budget cap as a circuit breaker).

Proposed solution: “TokenOps” platform (run-level governance)

The proposed platform includes:

  • Cumulative budgets across attributed runs, not just per-request caps.
  • Enforcement on the call path (run-level), enabling actions that steer agents before killing them.
  • An architecture with an out-of-band (outbound) control plane that does not interfere with application code.

Architecture (out-of-band plane) and modules

An outbound control plane is described with three major modules:

  1. Instrumentation

    • Observability using OpenTelemetry
    • Cost telemetry and enrichment
    • Attribution: identify which agent run / user dimension caused spend
  2. Accounting

    • Ledger-style accumulation of total runs/spend
  3. Enforced layer

    • Steering and policy execution
    • Halt as final protection when budgets are exhausted

Demo/code-level integration concepts

Key integration mechanism: a boundary annotation.

  • Applied to methods (intended to work across frameworks such as LangChain).
  • Tracks input/output
  • Forwards traces to the control plane
  • Writes ledger entries tied to run IDs and attributes
  • Provides a channel for the control plane to push actions back into agent behavior

Additional integration notes:

  • The control plane runs inside the tenant (claimed to reduce data leak concerns).
  • A governor component uses developer configs defining which actions are allowed (to prevent arbitrary or untrusted changes to agent behavior).

Control plane: segments, budgets, actions, policies

Inside the control plane:

  • Segmentation

    • Group runs by “dimensions” (e.g., cohort tags like user cohorts)
    • Apply budgets at different granularities (agent-level, run-level, cohort-level rollups)
  • Budgets

    • Static thresholds over time windows for segments/runs
  • Actions (two flavors)

    • Halt actions: kill/stop the agent if budget is exceeded
    • Steer actions: modify behavior to fit within budgets (avoid immediate termination)
  • Policies

    • Combine budgets + actions + conditions
    • Execute enforcement for targeted segments/runs

Policy catalog / failure modes covered (examples)

The approach benchmarks against policy coverage intended to address:

  • Spend management
    • budget/guardrails
  • Context management
    • context compaction
    • tool output reduction
  • Loop detection and progress detection
  • Additional run control mechanisms for “runaway token” failure modes

Steering examples (behavioral changes)

A “cost guard” mechanism can:

  • Consider both:
    • budget consumed so far, and
    • velocity (rate of token consumption)
  • If it predicts running out by the end of the run, it injects guidance into system instructions, e.g.:
    • “be more succinct”
    • “summarize more”

Goal: reduce future token usage rather than killing immediately.


Demo scenarios and results

Scenario 1: Preview mode

  • Policies are evaluated, but enforcement actions are not executed.
  • Dashboard shows which policies would run; run completes without failures.

Scenario 2: Governance enabled (halt behavior)

  • If the agent exceeds the cost cap, it gets killed immediately (circuit breaker).

Scenario 3: Steering behavior (cost guard)

  • Budget is higher than the halt scenario but still insufficient for normal completion.
  • Instead of killing, the system steers generation to fit the budget.

Benchmark claims

  • Tested on open source repos / stress tests (examples named: Browser-use and MetaGPT).
  • With the full policy suite enabled:
    • Average spend reduced ~78%
    • Completion rate improved from ~67% to ~96% compared to simple throttling (which kills runs)

Speaker / source identification

  • Tisha Chawla (presenter)
  • Susheem Koul / Sashim (co-presenter; demo presenter)
  • Source context: Microsoft event/talk (per video title)

Original video