Definition

LLM model routing (a.k.a. model routing / LLM routers) is the pattern of placing a classifier between the application and model providers that inspects each incoming request, judges its complexity or category, and dispatches it to whichever model promises the best cost/capability trade-off — “call a cheap model for simple tasks, a powerful one for hard ones”. A 2026 wave of router launches promised inference cost reduction as their core value.

The Counter-Position: One Router Vendor Deprecated Its Own Router

Manifest shipped an LLM router in March 2026 (four complexity tiers: simple, standard, complex, reasoning), deprecated it in June, and shut it down 2026-09-01 after four months across ~7,000 cloud users with mixed results. Its post-mortem is the most data-backed published critique of the pattern:

  1. Complexity cannot be deduced from the prompt alone. The prompt is a trigger, not the task; true complexity surfaces later through tool calls and web searches. The same instruction (“evaluate and improve the tests for $GIT_REPO”) spans trivial to massive depending on the repo.
  2. Caching beats routing for cost. Cache reads run 75–90% cheaper than uncached inputs, and the biggest token consumers (system prompts, conversation history) sit at the prompt start where prefix caching works best. A cache-aware router must add stickiness to the model it picked — doing its job by not doing it.
  3. Routers break behavior consistency. Engineers who choose models deliberately (the painter-knows-the-brush argument) produce better work than an opaque classifier switching models mid-session; unpredictability fragments evals, system prompts, and observability, and “managing that extra layer of uncertainty can cost more than it saves.”

The counter-claim stands too: routing advocates argue engineers shouldn’t have to track dozens of models, and routing remains plausible for high-volume, low-stakes, stateless workloads. Both positions are retained here (see contradictions).

  • llm-wiki — the artifact-first alternative: model choice as deliberate engineering decision documented per task, not delegated to a classifier
  • manifest — the gateway vendor that shipped, deprecated, and then pivoted away from LLM routing; also shipped the Autofix self-healing layer that now covers general API calls
  • rag — retrieval is another lever on the same cost/capability curve (pay retrieval cost, use a smaller model)

Implications

For agent harnesses (this vault’s operating context), the post-mortem favors: explicit per-request model choice, prefix caching on stable system prompts, and per-request isolation with tuned params — instead of on-the-fly complexity classification. If routing is ever worth revisiting, the burden of proof is a measured eval on this vault’s workload showing savings after stickiness and consistency costs.

Sources

^[raw/external/manifest-build-why-we-deprecated-our-llm-router-9e78a424.md]