Token Optimizer MCP
What It Is (confidence: medium)
Token Optimizer MCP (ooples/token-optimizer-mcp) is an MCP server + agent plugin that optimizes token usage for AI coding agents (Claude Code, Codex, Cursor, Windsurf, Cline, Gemini CLI, OpenCode, Copilot CLI). Claims “95%+ token reduction” via caching, compression, and smart tool intelligence. Ships a knowledge graph (“the part nothing else has”) built from the user’s own transcripts to measure prompt-cache economics, decide model routing by outcomes, and detect waste.
Snapshot 2026-09-01 confirms headline Measure token savings per AI coding agent, optimize context, and share a live local knowledge graph across 16 CLI clients — scope expanded to 16 clients and live graph sharing.
Distinctive Design (confidence: medium)
- Compaction is consolidation, not loss — frames compaction as lossless consolidation, unlike raw truncation.
- Progressive disclosure that tracks what was already asked, so repeated context is not re-sent.
- Verification against a control arm — “proves its own savings against a control arm,” and an “Honest comparison” section acknowledges what other optimizers do and do not do.
- Enforcing tier — hooks that can refuse a wasteful tool call (pre-execution veto) rather than merely advising.
- Prompt-cache economics measured from your own transcript — cache hits and misses computed from actual sessions, not estimates.
- Live local knowledge graph across 16 clients — shared context graph (new in snapshot) as differentiator vs pure proxy compression.
Claims (confidence: low)
The 95%+ headline reduction number is a marketing claim from the repo description; the tool’s own architecture stresses measurement from transcripts and honest comparison rather than a single headline figure. get_optimization_report MCP tool exposes token/compression stats on demand. Snapshot 2026-09-01 confirma Measure token savings per AI coding agent como headline atual, ainda sem benchmark independente publicado.
Contested: input-only proxies (RTK) have shown +18% billed cost via output bloat; cache is ~87% of cost per TRINCR — Token Optimizer’s control-arm measurement is the guardrail, headline 95% remains unverified.
Implications (confidence: medium)
Token Optimizer MCP occupies the “measured and enforced” end of the optimizer spectrum: it claims to justify savings against a control arm and refuses wasteful calls. This is the strongest guard against the RTK-style “input savings < output increase” failure, because it measures the whole transcript rather than compressing one command stream. The knowledge-graph/transcript approach is closer to a feedback loop than a static proxy. Live graph across 16 clients extends the loop to fleet level, similar to Bifrost but for optimization intelligence.
Implications: Prefira medição com control arm sobre headline — exija get_optimization_report por tarefa antes de aceitar 95%.
Refresh — snapshot 2026-09-03 (ea1b1b42)
Recapture 2026-09-03 confirms and extends the 09-01 state; headline now reads “Spend less context, keep the conclusions, and audit every claim across 16 coding clients” across 16 coding clients (~507★/61 forks at capture). COMPOUND adds (confidence: medium, vendor-sourced):
- Knowledge graph “why this is not RAG” — traversal + lexical search, not similarity: deterministic, instant, explainable, works offline; “retrieves verdicts — the reasoning already happened”.
- Zero-turn refusal — a plain deny costs a full turn (model calls Read, is refused, re-plans, then
smart_read), motivating enforcement that refuses at the right point. - Enforcing tier widened to 10 clients with pre-execution hooks (Claude Code native plugin, Codex, GitHub Copilot CLI, Gemini CLI, Qwen Code, Cursor, et al.). Plugin-first install (“install the plugin, not the bare MCP server”,
/plugin marketplace add ooples/token-optimizer-mcp). fleet_audit/ one audit across every project — ranks whole machine by measured cost; a fix proven in one project is offered to the others containing the same file contents, carrying evidence by content hash.- Waste detection as a ratchet — a report read once and forgotten becomes a persistent enforcement signal.
- Model routing decided by outcomes, not task size; cross-session memory holds “findings, decisions, dead ends”; compaction “consolidation, ranked by cost-to-re-derive”; cache economics “measured from the transcript, attributed to a line”.
- Verification suite ship —
verify:all: test 2,437 / verify:clients 253 / verify:harvest 26 / verify:ui 22 / doctor 10;npx token-optimizer-doctorfeeds a synthetic payload to the real hook binary. - Diagnostics — every native hook writes a bounded JSONL lifecycle record; raw event rows opt-in, capped at 1,000.
- MIT license (commercial use ok) — honest-comparison table flags typical alternatives as often noncommercial-only.
- Headline 95% reduction remains unverified; control-arm + honest-comparison framing persists as the guardrail. Star/fork counts are snapshot-time and volatile, not stable facts.
Open Questions
- Whether the “95%+” headline survives the tool’s own honest-comparison benchmark at real task mix.
- Whether the pre-execution veto hooks degrade developer trust / agent flow in practice.
- Does the 16-client live graph introduce privacy/cost overhead that offsets token savings?
Related
rtk-rust-token-killer — proxy-based input compression; the failure mode this tool measures against
vibe-coder-mcp — MCP server whose map-codebase tool claims 95-97% token reduction (README, unvalidated); same token-cost lever from inside the agent
caveman — complementary output-side reduction: /caveman skill cuts ~65% output tokens, caveman wrap is the local proxy sibling
headroom — compression proxy with output shaping
token-usage-reduction — the cost-lever space this tool automates
skill-reducer — complementary research on compressing authored skill content
claude-code-system-prompt — primary target agent (plugin install path)
bifrost — gateway peer with enterprise governance vs this per-agent optimizer
External Links
Sources
- raw/external/github-com-token-optimizer-mcp-a42670b5.md
- raw/external/github-com-token-optimizer-mcp-ea1b1b42.md