OAuth au niveau MCP : Comment nous avons résolu la gestion des jetons d'entreprise pour les agents d'IA

Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises
Une présentation détaillée de l'architecture d'authentification au sein de la passerelle MCP de TrueFoundry — trois plans, deux flux OAuth, un coffre-fort de justificatifs d'identité — racontée du point de vue d'un client qui la déploie en production.
Un lundi matin chez Northwind. Un lundi à 8h47, l'ingénieur d'astreinte de Northwind Logistics — une plateforme de chaîne d'approvisionnement fictive mais malheureusement plausible, comptant 50 ingénieurs — reçoit un ping Slack. Leur assistant IA interne, Cargo Copilot, renvoie des erreurs 401 pour chaque appel d'outil Jira et GitHub. La cause profonde est l'ordinateur portable d'un développeur, perdu à l'aéroport pendant le week-end. Le service informatique a révoqué la session SSO de l'ordinateur portable. Mais l'ordinateur portable contenait également des jetons OAuth pour GitHub, Slack, Jira et un service d'analyse interne, chacun copié dans les paramètres de l'éditeur et les dotfiles du développeur. La sécurité ne peut pas dire avec certitude quels scopes ont été exposés, sur quelles autres machines ces jetons ont également été dupliqués (parce que les développeurs s'entraident), ni comment les révoquer tous sans perturber la matinée de tout le monde.
L'incident de Northwind n'est pas un cas isolé. C'est à cela que ressemble la gestion des justificatifs d'identité dans la plupart des entreprises six mois après une adoption enthousiaste du MCP. Le reste de cet article est l'histoire architecturale de la façon dont Northwind — et d'autres clients de TrueFoundry — cessent de reproduire cet incident.
1. Le problème de la prolifération des justificatifs d'identité : À quoi ressemblent réellement 50 développeurs × 8 serveurs MCP
Northwind compte environ cinquante ingénieurs utilisant activement des IDE compatibles MCP avec huit serveurs — GitHub, Slack, Google Workspace, Jira, un service d'analyse interne, une API logistique interne, Zendesk et un service propriétaire de planification d'itinéraires. Cela représente quatre cents connexions actives selon un calcul simple, et en pratique, un nombre considérablement plus élevé de placements de justificatifs d'identité si l'on compte les entrées de trousseau, les paramètres d'éditeur, les fichiers d'environnement shell, les variables CI et les copies que les gens font lors de l'intégration d'un coéquipier.
Le rayon d'impact d'un poste de travail compromis n'est pas d'un seul jeton. C'est un jeton multiplié par les scopes attachés à ce fournisseur, la durée de vie du jeton (souvent 90 jours ou plus pour les jetons d'accès personnels), le nombre de machines qui l'ont dupliqué, et l'absence de révocation centrale. Si un jeton OAuth Slack permet la lecture de l'historique des canaux et qu'un jeton GitHub co-localisé permet la lecture des dépôts, une seule exfiltration devient un incident inter-systèmes avant que l'équipe de sécurité n'ait eu le temps de réagir. Les justificatifs d'identité MCP locaux sont également opérationnellement invisibles : pas de piste d'audit unifiée, pas de rotation forcée, pas de vérification de politique avant l'exécution d'un outil, pas de moyen rapide de savoir qui peut appeler quoi.
Le modèle de passerelle comble cette lacune. Les développeurs s'authentifient une seule fois auprès de la passerelle, et celle-ci gère l'authentification en aval selon le modèle configuré de chaque serveur — ce qui correspond exactement à la conception décrite dans les Présentation de la passerelle MCP de TrueFoundry et les documents d'authentification et de sécurité. L'avantage n'est pas seulement la consolidation. C'est que l'authentification devient une propriété de l'infrastructure plutôt qu'une propriété de chaque ordinateur portable de l'entreprise.
2. L'architecture d'authentification à trois plans : Entrant, Contrôle d'accès et Sortant
La première décision de conception est de séparer trois plans que les configurations de développeurs locaux confondent couramment. Les documents d'authentification et de sécurité de TrueFoundry décrivent cela comme un système en trois parties : authentification entrante, contrôle d'accès, et authentification sortante. La documentation stipule explicitement que les couches entrantes et sortantes sont indépendantes, ce qui constitue un avantage architectural majeur.
- Authentification entrante. Qui appelle la passerelle ? TrueFoundry prend en charge quatre méthodes : Clé API TrueFoundry (PAT), Jeton de compte virtuel, Jeton de fournisseur d'identité (un JWT Okta/Auth0/Azure AD validé via un fournisseur d'identitéconfiguré), et OAuth TrueFoundry (utilisé par des IDE comme Cursor et VS Code).
- Contrôle d'accès. Que peut faire l'identité résolue ? Chaque serveur MCP dispose de Collaborateurs — utilisateurs, équipes ou comptes virtuels — avec des permissions basées sur les rôles. La portée au niveau de l'outil est obtenue en combinant des outils de plusieurs serveurs MCP en un Serveur MCP Virtuel, qui n'expose qu'un sous-ensemble sélectionné.
- Authentification sortante. Comment la passerelle atteint-elle chaque serveur en aval ? Sept méthodes sont prises en charge : OAuth2 (Code d'autorisation), OAuth2 (Identifiants client), Clé API (Partagée), Clé API (Individuelle), Aucune authentification, Transmission de jeton, et Redirection de jeton via x-tfy-mcp-headers.

Le principe de conception est simple : l'identifiant utilisé pour atteindre la passerelle n'est jamais l'identifiant qui atteint l'outil en aval. La rotation entrante n'affecte aucun état OAuth du fournisseur. Un serveur MCP GitHub peut utiliser OAuth par utilisateur tandis que le service d'analyse interne utilise les identifiants client — dans le même agent, dans le même flux de requête. La passerelle est à la fois le point d'application de la politique et le courtier d'identifiants.
3. 2LO vs 3LO : Choisir le bon type d'octroi OAuth pour chaque serveur MCP
Dans le jargon d'entreprise, 2LO et 3LO désignent deux types d'octroi de RFC 6749 (OAuth 2.0), mis en correspondance avec la manière dont les utilisateurs finaux participent.
OAuth à deux pattes (2LO) utilise le flux d'octroi des identifiants client . La passerelle détient un client_id et client_secret, envoie grant_type=client_credentials au point de terminaison de jeton du fournisseur, reçoit un access_tokende courte durée, et utilise ce jeton à chaque requête vers le serveur MCP. Pas de navigateur, pas de consentement utilisateur. C'est le bon choix pour les intégrations de serveur à serveur : API d'analyse internes, microservices backend, services de données internes.
OAuth à trois pattes (3LO) utilise le Code d'autorisation . Chaque utilisateur final est redirigé vers le fournisseur (GitHub, Slack, Google) pour autoriser l'accès à ses ressources spécifiques. La passerelle stocke un refresh_tokenpar utilisateur, maintient une correspondance chiffrée de user_id → {access_token, expires_at, refresh_token, scopes, provider_metadata}, et renouvelle automatiquement le jeton d'accès.
Comment Northwind a fait son choix
Les huit serveurs MCP de Northwind se sont répartis à peu près équitablement. GitHub, Slack, Jira, Google Workspace et Zendesk sont tous des outils par utilisateur — la valeur de « exécuter ceci en tant qu'Alice » l'emporte sur la surcharge opérationnelle d'un jeton par utilisateur. Ceux-ci utilisent le 3LO. Le service d'analyse interne, l'API Logistique et le service de planification d'itinéraires exposent tous des données à l'échelle de l'entreprise ; aucun OAuth par utilisateur n'a de sens, et une seule subvention de Client Credentials partagée est à la fois plus simple et plus facile à auditer. Ceux-ci utilisent le 2LO.
À quoi ressemble le modèle de configuration
Remarque sur cette vue de configuration
TrueFoundry enregistre les serveurs MCP via l'interface utilisateur (Paramètres → Serveurs MCP → Ajouter un serveur MCP) plutôt que via un fichier YAML géré par le client, bien que la configuration puisse également être appliquée de manière déclarative via tfy apply. Le fichier YAML ci-dessous présente le modèle de configuration par serveur — les mêmes champs que vous remplissez dans l'interface utilisateur — en utilisant le schéma vérifié de TrueFoundry (provider-account/mcp-server-group avec des intégrations).
# Northwind's backend-group MCP Server Group: two MCP servers, one per OAuth flow.
# Schema matches TrueFoundry's provider-account / integrations pattern (tfy apply).
name: backend-group
type: provider-account/mcp-server-group
collaborators:
- subject: team:platform-engineering
role_id: mcp-server-manager
- subject: team:site-reliability
role_id: mcp-server-manager
- subject: virtualaccount:cargo-copilot-runtime
role_id: mcp-server-manager
integrations:
# ---- 3LO (Authorization Code): per-user OAuth on GitHub ----
- name: github
type: integration/mcp-server/remote
description: GitHub MCP server using per-user OAuth Authorization Code flow.
url: https://github-mcp.example.com/mcp
transport: streamable-http
authorized_subjects:
- team:platform-engineering
- team:site-reliability
- virtualaccount:cargo-copilot-runtime
auth_data:
type: oauth2
grant_type: authorization_code
authorization_url: https://github.com/login/oauth/authorize
token_url: https://github.com/login/oauth/access_token
client_id: tfy-secret://northwind:github-oauth:client_id
client_secret: tfy-secret://northwind:github-oauth:client_secret
scopes:
- repo
- read:org
jwt_source: access_token
code_challenge_methods_supported:
- S256
# ---- 2LO (Client Credentials): shared service identity ----
- name: internal-analytics
type: integration/mcp-server/remote
description: Internal analytics MCP server using OAuth2 Client Credentials.
url: https://analytics-mcp.internal.example.com/mcp
transport: streamable-http
authorized_subjects:
- team:platform-engineering
- team:site-reliability
auth_data:
type: oauth2
grant_type: client_credentials
token_url: https://northwind.okta.com/oauth2/default/v1/token
client_id: tfy-secret://northwind:analytics-oauth:client_id
client_secret: tfy-secret://northwind:analytics-oauth:client_secret
scopes:
- analytics:read
jwt_source: access_tokenLa commande pour appliquer le fichier YAML ci-dessus est :
tfy apply -f backend-group.yaml --dry-run --show-diffTrois choses à noter. Premièrement, tfy-secret:// est le schéma de référence FQN canonique que TrueFoundry utilise pour les secrets stockés dans un groupe de secrets TrueFoundry, avec le format séparé par des deux-points tfy-secret://<tenant>:<secret-group>:<secret-key> documenté dans le guide sur les variables d'environnement et les secrets. La valeur réside dans votre gestionnaire de secrets (AWS SSM, GCP Secret Manager, HashiCorp Vault, Azure Key Vault) ; TrueFoundry résout le FQN à l'exécution, de sorte que le fichier ci-dessus est sûr dans un dépôt Git. Deuxièmement, authorized_subjects est la liste RBAC par serveur — un mélange d'équipes et de comptes virtuels — ce qui permet au même serveur MCP de servir à la fois les développeurs humains (dont les PATs se résolvent en appartenance à une équipe) et les tâches planifiées de comptes de service (dont les jetons de compte virtuel sont mappés à virtualaccount: sujets) sans dérive de configuration. Troisièmement, l'entrée 3LO définit code_challenge_methods_supported: [S256] — PKCE est recommandé pour chaque client de type "Authorization Code" aujourd'hui et obligatoire dans la nouvelle spécification MCP du 25/11/2025.
Côté client, l'agent n'énumère jamais les modes d'authentification. Il compose une URL de passerelle, présente une seule information d'identification entrante, et la passerelle injecte la bonne information d'identification sortante par serveur :
from fastmcp import Client
from fastmcp.client.transports import StreamableHttpTransport
# Real URL pattern from the TrueFoundry docs.
transport = StreamableHttpTransport(
url="https://llm-gateway.truefoundry.com/mcp-server/backend-group/github/server",
headers={"Authorization": f"Bearer {user_token}"},
)
async with Client(transport) as client:
result = await client.call_tool(
"search_issues",
{"query": "repo:northwind/logistics-core is:open label:critical"},
)4. Gestion du cycle de vie des jetons : Actualisation proactive et accès concurrentiel
La correction OAuth ne se limite pas à l'obtention de jetons. Il s'agit de les utiliser en toute sécurité sous une charge concurrente. La conception naïve actualise lorsque le fournisseur renvoie 401, ou lorsque maintenant ≥ expires_at. Cette conception fonctionne en démonstration et échoue le premier lundi où elle rencontre la réalité.
Rejouons l'incident Northwind avec une actualisation naïve : le lundi à 8h47, cinquante ingénieurs se connectent, chaque travailleur découvre indépendamment le même jeton GitHub expiré, cinquante appels parallèles POST /token frappent GitHub depuis la même passerelle, et GitHub limite le débit de l'ensemble. Cargo Copilot renvoie des 401 à tout le monde jusqu'à ce que la file d'attente se vide.
Le modèle qui tient en production a deux ingrédients : actualiser de manière proactive à 80 % du TTL, et sérialiser l'actualisation par (utilisateur, serveur) avec un verrou distribué.
def get_token(user_id, server_id):
token = cache.get(user_id, server_id)
if now() < token.issued_at + 0.8 * token.expires_in:
return token
# Crossed the 80% threshold. Serialize refresh so 50 workers
# do not all hit the provider token endpoint at once.
with distributed_lock(user_id, server_id):
token = cache.get(user_id, server_id) # re-read under lock
if now() < token.issued_at + 0.8 * token.expires_in:
return token # someone else already refreshed
new = provider.refresh(token.refresh_token)
vault.put_atomic(user_id, server_id, new) # access + refresh together
cache.set(user_id, server_id, new)
return newL'actualisation à 80 % du TTL laisse une marge pour le décalage d'horloge, les tentatives, les points de terminaison de jetons lents des fournisseurs et les requêtes déjà en cours. Le verrou par (utilisateur, serveur) empêche un effet de "thundering herd" contre le fournisseur. La lecture à double vérification à l'intérieur du verrou n'est pas cosmétique : elle permet à la première requête d'effectuer l'actualisation tandis que les autres réutilisent le jeton fraîchement écrit.
Deux subtilités sont importantes à grande échelle. Premièrement, pour les intégrations de type "Client Credentials", la clé de cache est (server_id, provider_account); pour les intégrations de type "Authorization Code", la clé doit inclure user_id. Deuxièmement, les fournisseurs qui renouvellent les jetons d'actualisation à chaque actualisation — Google et Microsoft le font, selon la configuration — exigent que l'écriture soit atomique. Si le nouvel access_token is committed before the new refresh_token, a crash leaves the user stranded with an invalidated old refresh token and a forced re-authentication.
5. The Double-Auth Flow: A Complete Sequence Diagram
With the planes separated and the refresh path defined, an end-to-end MCP request through the gateway has a precise shape. Trace one Cargo Copilot call — "summarize the open critical-severity issues in repo northwind/logistics-core" — from agent to GitHub and back:

Three properties of this flow are worth pointing out. The agent never receives the GitHub access token; only the gateway holds it. The MCP server never sees Northwind's inbound TrueFoundry credential unless Token Passthrough is explicitly configured (more on that in §7). And the Northwind security team gets a single governed request path with identity, authorization, and audit recorded in one place — the same property that would have let them answer the Monday-morning question, "which scopes were on that laptop?", in under a minute.
The MCP authorization specification reinforces one part of this picture: bearer access tokens must be sent in the Authorization request header and never in a URI query string. That norm matters precisely because it keeps tokens out of access logs, browser history, and proxy traces.
6. Credential Vault Integration: Why Raw Secrets Never Touch Application Memory
A production gateway should not treat OAuth refresh tokens like ordinary application rows. Client secrets and refresh tokens are long-lived keys to the kingdom — leaking a Slack refresh token is, for most purposes, equivalent to leaking the Slack workspace.
The standard pattern is envelope encryption. The gateway generates a random data encryption key (DEK) for each token record, encrypts the OAuth secret with that DEK, encrypts the DEK with a master key in KMS or a secret manager, and stores the encrypted DEK alongside the ciphertext. On read, the gateway asks the vault to decrypt only the DEK, then decrypts the token in a narrowly scoped code path that avoids logging or serializing plaintext.
In stricter deployments, token exchange itself can be pushed into a sidecar or a vault transit operation, so application code never receives raw client_secret material at all. Even with vanilla envelope encryption the blast radius changes substantially: a database dump alone is insufficient (the DEKs are not stored decrypted), and a memory dump of a general request worker is much less likely to contain broad credential material (only the keys recently requested are decrypted, and only for the duration of the request).
Northwind's deployment uses AWS KMS as the master key store and the TrueFoundry secrets store for the encrypted material. The gateway never reads a provider client_secret directly; it reads a TrueFoundry secret FQN, fetches the ciphertext, and asks KMS to unwrap. The deployment-specific work that remains — KMS key rotation policy, audit retention, operational runbooks — is the responsibility of every enterprise running the gateway, and worth a dedicated review.
7. Token Passthrough vs. Token Forwarding: Edge Cases and When to Use Each
Two patterns sound similar and have very different security implications. They are worth distinguishing carefully because misusing either one undoes most of the gateway's value.
Token Passthrough is appropriate for internal MCP servers that validate the same enterprise identity token the gateway already validated inbound. Northwind's Logistics API is one example: the gateway and the API both trust the company's Okta tenant, so passing through the Okta JWT keeps the caller's identity intact end-to-end. The risk is token audience confusion: a token minted for the gateway should not automatically be valid for every MCP server. RFC 8707 (Resource Indicators) exists precisely to bind tokens to an intended resource, and the MCP authorization spec now requires that clients include a resource parameter and that servers validate that presented tokens were issued for them.
Token Forwarding works differently. TrueFoundry's x-tfy-mcp-headers header carries a stringified JSON keyed by the remote MCP server identifier in <group>/<server> format. Here is the format verbatim from the Virtual MCP Server docs:
import json
from fastmcp import Client
from fastmcp.client.transports import StreamableHttpTransport
extra_headers = json.dumps({
# Custom auth for the backend-group/sentry MCP server backing this
# virtual server. Overrides whatever outbound auth was configured.
"backend-group/sentry": {"Authorization": "Bearer your-sentry-token"},
})
transport = StreamableHttpTransport(
url="https://llm-gateway.truefoundry.com/mcp-server/backend-group/restricted-sentry/server",
headers={
"Authorization": f"Bearer {tfy_token}", # inbound auth (gateway)
"x-tfy-mcp-headers": extra_headers, # outbound override
},
)Treat Token Forwarding as an emergency hatch, not a default. The gateway's value is exactly the token management it skips.
8. Latency Benchmarks: What Token Management Adds to Every Tool Call
A gateway adds work before the downstream MCP call: inbound auth validation, policy lookup, token lookup, and occasionally a refresh. In a healthy deployment the common path is a cache hit — token metadata is valid, the encrypted token reference is warm in memory.
Scenario envelopes
Combining the component costs above, the gateway adds roughly:
Two takeaways. First, the common path is a cache hit, and at single-digit milliseconds it is invisible against the surrounding LLM call (typically hundreds of milliseconds to seconds) and the downstream MCP server (typically tens to hundreds of milliseconds). Second, the cache-miss-plus-refresh path is rare but expensive enough to matter when it stacks up under load — which is exactly why §4 specifies proactive refresh at 80% of TTL with a per-(user, server) lock. With proactive refresh, real users almost never see the refresh latency on the request path; the gateway has already prefetched a new token.
Every team should still instrument server-timing headers, token endpoint duration, lock wait time, vault decrypt time, and downstream MCP latency separately. The goal is to prove that token management is not the dominant cost in the agent's end-to-end tool call — the envelopes above suggest it almost never is, but you should verify with your own provider mix.
The operational fix for the cache-miss case is background refresh. When the gateway sees a token cross the 80% threshold, it refreshes on the request path once, then schedules subsequent refreshes before users notice. For high-volume shared Client Credentials integrations this almost eliminates user-visible refresh latency. For per-user Authorization Code integrations, the system can prioritize active users and frequently used MCP servers.
Before vs. after: security posture at Northwind
9. FAQs
Is OAuth at the MCP layer mandatory?
No. The MCP authorization specification makes authorization optional, but HTTP-based implementations that support it should conform to the spec. In practice, enterprises with non-trivial tool estates should centralize auth long before they consider it strictly mandatory.
Should every MCP server use 3LO?
No. Use Authorization Code (3LO) when the agent acts on a human user's resources — Slack, GitHub, Gmail, CRM tools. Use Client Credentials (2LO) when the downstream server authenticates the application or gateway as itself. Mixing both within one Virtual MCP Server is normal and expected.
Does a gateway eliminate provider-side permissions?
No. It complements them. Slack, GitHub, Google, and internal APIs still enforce their own scopes and resource permissions. The gateway adds a coarser, faster layer of policy on top — server-level, tool-level, team-level — and a uniform audit trail across all of them.
OAuth 2.0 or OAuth 2.1?
The TrueFoundry gateway speaks the OAuth 2.0 grants real-world providers support today: Authorization Code (with PKCE), Client Credentials, and refresh tokens. The MCP authorization specification itself targets OAuth 2.1, which is largely OAuth 2.0 plus the security best practices that are already considered mandatory in modern deployments (PKCE everywhere, no implicit grant, no password grant). For practical purposes the two are convergent; we follow OAuth 2.1 guidance where it tightens behavior.
What about the newer 2025-11-25 MCP spec?
The November 2025 revision adds mandatory PKCE for all clients, formalizes Client ID Metadata Documents (CIMD) as the preferred client registration method, and introduces step-up authorization for incremental scope consent. These tighten security without invalidating the architecture in this post. We are tracking them and rolling support in as it matures across the ecosystem.
Where does TrueFoundry fit?
The TrueFoundry MCP Gateway is a production implementation of this architecture: a centralized server registry, configurable inbound auth, access control, multiple outbound auth modes, and token lifecycle management. This post describes the design; the Authentication and Security docs are the operational reference. For the implementation patterns covered here, see also the docs on Auth Overrides and Virtual MCP Server.
Take the next step
If you are a platform, security, or AI infrastructure leader running MCP at non-trivial scale, the highest-leverage action is to inventory every MCP server in active use, map each one to the correct inbound, policy, and outbound auth model, and decide where centralization actually wins. Northwind's mapping took two afternoons; the migration itself took a single sprint. We are happy to walk through that exercise with your team.
Read the architecture: TrueFoundry MCP Gateway authentication and security. Or book an enterprise security architecture review with our team.
Further reading
All citations in this post are linked inline. The references below collect the same URLs for printability and link-rot insurance.
- TrueFoundry MCP Gateway overview. https://www.truefoundry.com/docs/ai-gateway/mcp/mcp-overview
- TrueFoundry MCP Gateway authentication and security. https://www.truefoundry.com/docs/ai-gateway/mcp/mcp-gateway-auth-security
- TrueFoundry Auth Overrides. https://www.truefoundry.com/docs/ai-gateway/mcp/mcp-server-auth-overrides
- TrueFoundry Virtual MCP Server. https://www.truefoundry.com/docs/ai-gateway/mcp/virtual-mcp-server
- TrueFoundry using MCP Gateway in your agent. https://www.truefoundry.com/docs/ai-gateway/mcp/use-mcp-gateway-in-agent
- TrueFoundry Identity Providers. https://www.truefoundry.com/docs/platform/identity-providers
- TrueFoundry Secret Manager in AI Gateway. https://www.truefoundry.com/docs/ai-gateway/secret-manager-in-ai-gateway
- TrueFoundry Environment Variables and Secrets (tfy-secret:// FQN format). https://www.truefoundry.com/docs/environment-variables-and-secrets
- TrueFoundry Using tfy apply (provider-account / integrations YAML pattern). https://docs.truefoundry.com/docs/using-tfy-apply
- Okta Concurrency Limits (API response time guidance). https://developer.okta.com/docs/reference/rl2-concurrency/
- WorkOS: RS256 vs HS256 (JWT verification benchmarks). https://workos.com/blog/rs256-vs-hs256-jwt-signing-algorithms
- AWS KMS documentation (decrypt latency guidance). https://aws.amazon.com/kms/faqs/
- Model Context Protocol authorization specification, 2025-06-18. https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization
- RFC 6749: The OAuth 2.0 Authorization Framework. https://datatracker.ietf.org/doc/html/rfc6749
- RFC 8707: Resource Indicators for OAuth 2.0. https://datatracker.ietf.org/doc/rfc8707/
- RFC 9728: OAuth 2.0 Protected Resource Metadata. https://datatracker.ietf.org/doc/html/rfc9728/
Remarque : Northwind Logistics est une entreprise fictive utilisée pour ancrer la conception dans un déploiement concret. Le modèle de configuration présenté au §3 est illustratif ; les serveurs MCP de production sont enregistrés via l'interface utilisateur de TrueFoundry ou appliqués via tfy apply.
TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.













.webp)



.png)
.png)
.png)
.png)
.png)






.png)







