Trimmus sits between your app and OpenAI or Anthropic. It audits every call, optimizes the ones that repeat, and cuts your costs without losing quality. The same answers, for less.
Point your app at Trimmus instead of the provider. Change the base_url, use your Trimmus key. Your code, SDKs and streaming stay exactly as they are.
Every call is logged, priced, and attributed to the feature behind it. You see what each feature costs to run, in dollars, not vague percentages.
Trimmus removes spending you did not need, automatically, on every request. It never changes what your model returns. If a request cannot be safely optimized, it passes through untouched.
Every saved dollar is tied to a specific request you can look up. No blended estimates.
Monthly totals are reconciled with what the provider actually charged. If the math drifted, the report would say so.
Trimmus never trades quality for savings. Your prompts are not rewritten and your responses are not altered.
An estimate from typical traffic patterns. Your real number comes from a pilot on your real traffic.
No percentage of savings, no usage markup. One subscription, sized to your monthly spend. Built for teams spending $2k or more a month.
Every plan starts with a two-week pilot on your real traffic. The savings report decides if we’re worth it.
Before a team routes real traffic through us, these are the questions that come up. Straight answers.
No. Trimmus adds under 10ms to a call. Many optimized requests come back faster, because a cache hit skips the provider round trip entirely.
Your traffic keeps flowing. Trimmus is fail-open by mandate. If any optimization errors or the service is unavailable, the original request passes straight through to the provider. Our bug is never your outage.
No. Only zero-loss techniques ship: exact caching, cache structuring, and request coalescing. No compression, no model downgrades, no semantic rewriting. The response is byte-for-byte what the provider would have returned.
You keep your own provider keys and authenticate to Trimmus with a Trimmus key. Prompts and responses are used only to serve your cache and build your audit, never to train models. Retention is configurable.
Most AI bills repeat themselves. Identical and overlapping requests are served from an exact cache, prompt prefixes are structured to hit the provider cache, and concurrent duplicate calls are coalesced. You stop paying twice for the same tokens.
Every saved dollar is tied to a request you can open: cost with Trimmus, cost without, and the difference. Monthly totals are reconciled against the provider invoice. If the math drifted, the report says so.
Start with a two-week pilot on your real traffic. The audit shows your actual number before you commit. If the savings are not there, you walk away.
Two weeks on your real traffic. A full audit of what every feature costs. If the savings are not there, walk away.
We reply within one business day.