Video summary

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

Main summary

Key takeaways

Technology

Main topic

The talk argues that model routing should be driven by request-specific “preferences” (cost/latency/quality/rules/task context) rather than by always selecting the top benchmark model. It positions model routing as a key capability in DigitalOcean’s inference engine.

Why “one best model” is the wrong approach (3 reasons)

  1. Cost explosion
    • Inference spending is rapidly increasing; even large companies (examples mentioned: Walmart, Uber, Microsoft) cap usage to control inference bills.
  2. Fit / overkill
    • Using one frontier / best-in-benchmark model for every task is usually inefficient because many tasks can be handled by smaller or cheaper models.
  3. Risk / reliability
    • Relying on a single model creates a failure mode (no failover if the model degrades/outages).
    • Model orchestration is presented as the “new phase” for production robustness.

Core premise: there is no single “best model,” only the right model per request

The “right model” depends on factors that benchmarks can’t fully capture:

  • Task/request type (e.g., classification vs. code completion vs. code review/security)
  • System prompts and tools used with the model
  • Cost budget and willingness to trade off
  • Latency requirements (not all use cases need the same responsiveness)
  • End-user preference

Product/architecture: DigitalOcean Inference Router (live demo)

Key issues with prior “auto routing” approaches

  • Earlier attempts felt like a black box; if routing produced poor results, builders couldn’t easily improve it.

How this router is designed differently

  • Requests go through:
    • an Open Proxy and
    • a purpose-built routing model (both stated as open source)
  • No vendor lock-in emphasized (open components; ability to keep control).

Router customization inputs

Users provide what matters for their workload:

  • cost, latency, quality
  • preferred models
  • hard rules
  • task description in natural language

Router execution includes:

  • presets
  • rule/decision-tree style controls
  • ability to change options “in a single line of code”

Users validate using their own evaluations, not just public leaderboards.

Routing model performance claims

  • < 200 ms routing decision time
  • “Costs customers nothing extra” (router itself is free/included)
  • In evaluations, routing is claimed to outperform frontier models on routing tasks with fractional latency.

Features highlighted in the demo

UI configuration

Demonstrated router presets such as:

  • Software engineering
  • General writing
  • Knowledge bases
  • Document intelligence

Example custom “Software engineering” tasks:

  • bug fixing
  • code generation
  • test writing

Multi-model pools per task were shown.

Two routing strategies shown

  • Manual ranking / preferred model with failover
    • Example: for code generation, always try GLM 5.2; if it fails, fail over to GPT 5.2
  • Fastest/recency-based selection
    • Example: for bug fixing, choose the model that’s been fastest in the last ~30 minutes from a pool.

Demo results (what was measured)

Playground side-by-side tests

  • Compared routing vs. always using a single premium model (e.g., “Opus” mentioned).
  • Examples of tasks:
    • Fibonacci function → router chooses cheaper/faster model based on code-snippet intent
    • “Optimize my function” → routes to GPT 5.2
    • “Write some unit tests” → routes to a model for test writing / code verification
  • The pattern: faster + cheaper while still maintaining quality.

Evaluations

  • An evaluation compared:
    • Opus vs. routed approach
  • Reported that correctness/quality scores were close (within judge margin of error), while routing used fewer tokens and was significantly faster.

Real workflow / observability with OpenCode

  • Two terminals compared:
    • Left: single-model approach (always “Opus”)
    • Right: OpenCode configured to use the software engineering router
  • Live observability shown:
    • token usage in real time
    • which models were selected
    • task-to-model mappings
    • cost accumulating live
    • latency per step

Across session steps, reported large cost reductions, e.g.:

  • One session step: router ~8 cents vs. Opus ~25 cents (~3x cheaper)
  • Another: total router 14 vs Opus 44 (units implied as cents/total session cost)

Claimed latency is improved because routing selects appropriate models per step.

“Routing is a foundation” — what’s built on top

The talk lists three higher-level layers that leverage routing:

  1. Eval: validate the right model for your specific use case with your tests
  2. Caching: avoid paying repeatedly for the same answers
  3. Personalization: router learns what works for your team over time

It emphasizes a continuous improvement loop: route → evaluate → adjust.

Quick factual claims (explicit “quick facts” section)

  • Routing decision time: under 200 ms per request
  • 0 application code changes needed to adopt
  • Free/included (no separate router cost or rollout burden)
  • Router model is open sourced (referred to as via “Plano”)
  • Router sits in the inference engine layer (DigitalOcean stack context described earlier)

Main speakers / sources

  • Archana (Archa) Kamath — VP of Engineering for Inference Engine & AI Infrastructure, DigitalOcean
  • Tyler Gillam — built parts of the router; performed the live demo

Original video