An LLM gateway that pays for itself
Routing, caching and per-team budgets in front of three model providers, built in a fortnight, mostly with things we already ran.
When I started this rollout, every team had its own API key and its own idea of what a reasonable spend looked like. I built the gateway so that cost and quality decisions happen once, in one place.
Semantic caching on embeddings knocked 22% off request volume. Cheap-model-first routing with an escalation rule handled another 31%.
Budgets are enforced at the namespace level and surface in the same cost dashboard as compute, which is the only reason anyone looks at them.
Before the gateway, ten teams shipped prompts independently with no shared guardrails. Some retried aggressively on timeouts, some sent full chat history on every request, and none had a consistent way to cap monthly usage by team. Finance saw one bill. Engineering saw ten disconnected systems.
We started with three controls: request normalization, model routing policy, and per-namespace budget checks. Normalization trimmed duplicated system messages and limited conversation windows by endpoint type. Routing defaulted to lower-cost models and escalated only when a confidence threshold was missed.
Caching worked only after we were strict about keys. Naive text hashing missed equivalent requests with tiny formatting changes, so hit rate was poor. We moved to semantic keys on normalized input and TTL by use case, which pushed cache hit rate above thirty percent on internal support workloads.
The hardest part was not code. It was ownership. I asked each team to nominate one reviewer to approve prompt changes that increase token budgets, plus one shared review for model-policy changes. That slowed random experimentation but removed the "who changed this" loop during incidents.
After six weeks, cost variance flattened and response latency improved because retry storms dropped. The gateway did not make models smarter. It made behavior predictable, measurable, and affordable under real traffic.