There’s a line in your codebase that reads something like model=”gpt-4o”. Probably in several places. It was the obvious thing to write, and it quietly couples four unrelated concerns to your deploy pipeline: which provider serves the request, what happens when that provider is down, who is allowed to spend money on it, and what gets checked before the prompt leaves your network.
The fix isn’t a better string. It’s a layer.
What breaks at scale
Provider outages are your outages. Look at any major provider’s status page over a six-month window and you’ll find multiple incidents and stretches of degraded performance. If your application names one provider directly, each of those is your incident too, and your mitigation is a code change under pressure.
Latency varies enormously. Not just across models — across regions, across times of day, across providers for the same nominal model. Static routing means you’re regularly paying more latency than you need to, and you can’t see it, because you have nothing to compare against.
Rate limits arrive without warning. Provider quotas are enforced per account and per model. A traffic spike, or a colleague’s batch job on the same account, and your production agent starts getting 429s. Retrying against the same endpoint doesn’t help.
Cost is unattributable. The provider bill tells you what you spent. It doesn’t tell you which team, which application, or which user. Without attribution you cannot set a budget, so cost control becomes a quarterly conversation about the total rather than a control that prevents the problem.
Credentials sprawl. Every service that calls a model needs a key. Keys land in env vars, CI secrets, someone’s local .env, an agent definition. Rotation becomes an archaeology project.
The indirection that fixes most of it
Call a logical name instead of a real model. production-chat, not openai/gpt-4o. Behind that name, configure a routing strategy and one or more real targets — azure/gpt-4o, openai/gpt-4o, anthropic/claude-*.
TrueFoundry calls these virtual models, and the concept appears under various names across gateway products. The important part is what becomes possible once your application no longer knows which provider serves it:
Health-aware routing. A target returning errors gets taken out of rotation until it recovers. Your application doesn’t retry — the gateway does, against a different target, within the same request.
Weighted and latency-based distribution. Split traffic across providers by weight to stay under per-account rate limits, or route dynamically to whichever target is currently fastest.
Priority fallbacks. Primary target, then secondary, then tertiary. A provider outage becomes elevated latency rather than downtime.
Zero-deploy model swaps. New model ships, you point the logical name at it, you roll it back if quality regresses. No release, no code review, no coordination across six services that all hard-coded the old name.
One caveat worth knowing: this pattern generally applies to synchronous calls. Batch APIs run asynchronously on a single provider, so there’s no opportunity to fail over mid-batch — batch jobs typically have to name a real model directly.
Cost control that’s actually a control
Attribution first: every request should carry which user, team, application, and — if you’re running agents — which agent made it. Without that tag, everything downstream is reporting rather than control.
With it, you get budgets — a hard spend cap per user, team, or model that rejects requests when exceeded — and rate limits per user, per model, per application. The distinction matters: a rate limit prevents one client from starving others, a budget prevents a runaway loop from costing five figures overnight.
A runaway agent is the specific scenario to design for. An agent in a delegation loop can generate thousands of model calls in minutes. Rate limits slow that down; budgets stop it. You want both.
Semantic caching is the third piece and the one with the best payoff-to-effort ratio in practice. Repeat and near-repeat requests are far more common than people assume, particularly for RAG-style workloads where many users ask the same question differently. Cache on embedding similarity rather than exact match and you cut both cost and latency on the same requests.
Set budgets on the logical model name, not the real one — otherwise a fallback to a second provider routes around your spend cap, which is exactly backwards.
Guardrails belong at the boundary
Once applications handle real user data, and once agents call external tools on their own, the failure modes get concrete: a support bot echoing a credit card number because PII wasn’t stripped; a coding agent running a destructive shell command through an MCP tool; an internal Q&A bot jailbroken via prompt injection into leaking confidential data.
Guardrails inspect and, where needed, block or rewrite data before it causes damage. Four hooks are worth having: prompt before the model sees it, response before the user sees it, tool arguments before the tool runs, tool results after they return. The last two are agent-specific and the ones most often missing.
Two design decisions determine whether guardrails survive contact with production.
Operation mode. Validate inspects and blocks without changing anything — a moderation check that rejects a prompt outright. Mutate rewrites and can also block — a PII filter that turns “My SSN is 123-45-6789” into “My SSN is REDACTED” and lets the request through. Validation on input can often run in parallel with the model call, so it costs little latency; mutation has to run sequentially, in the request path.
Enforcement strategy. This is the one teams get wrong. What happens when a guardrail catches a violation is the easy half. What happens when the guardrail itself fails — times out, or its provider is down — is the half that causes outages.
Three positions: Enforce blocks on violation and blocks on guardrail error. Enforce but ignore on error blocks on violation, lets through on guardrail error. Audit logs everything and blocks nothing.
The rollout that works: start in Audit to see what would be caught without affecting users. Move to Enforce-but-ignore-on-error for real protection without a guardrail provider becoming a single point of failure. Move to strict Enforce only where compliance genuinely requires it. TrueFoundry’s guardrails documentation lays out this progression along with the parallel-versus-sequential execution detail, which is the part that determines your latency budget.
Going straight to strict Enforce is how teams end up with a guardrail vendor outage taking down their production chatbot — and then disabling guardrails entirely, which is worse than either option.
What you get for free
The underrated benefit of a single proxy is that observability stops being a project. Every request passes through one place, so cost, tokens, latency, error rates by status code, cache hit rates, and guardrail outcomes are recorded as a side effect. No SDK in every service, no coverage gaps from the team that forgot, and complete coverage of requests that were refused — which application-side instrumentation usually misses entirely.
If your organization already runs an observability stack, the gateway should export rather than replace: OpenTelemetry for traces and metrics, Prometheus scrape for dashboards and alerts.
The migration
The reason to do this is that it’s genuinely incremental.
- Put the gateway in the path for one non-critical service. OpenAI-compatible endpoints mean this is usually a base-URL change.
- Turn on logging and attribution. Two weeks of data will tell you things about your traffic distribution you didn’t know.
- Introduce logical model names, one per use case, each initially pointing at the single provider you’re already using. No behavior change.
- Add a fallback target to your highest-traffic name. This is the step that eliminates outage exposure.
- Add budgets and rate limits once attribution data shows you where the money goes.
- Add guardrails in Audit mode, then escalate per the progression above.
Steps 1–3 are a day’s work and change nothing observable. Step 4 is where the reliability payoff lands. Most teams stall between 3 and 4 because nothing is broken yet — which is exactly the right time to do it.
