Blank white background with no objects or features visible.

Nous vous offrons un accès gratuit à l'intégralité du Gartner Hype Cycle for AI Governance 2026. Obtenez votre exemplaire →

OAuth au niveau MCP : Comment nous avons résolu la gestion des jetons d'entreprise pour les agents d'IA

Par Boyu Wang

Published: October 10, 2026

Une présentation détaillée de l'architecture d'authentification au sein de la passerelle MCP de TrueFoundry — trois plans, deux flux OAuth, un coffre-fort de justificatifs d'identité — racontée du point de vue d'un client qui la déploie en production.

Key Takeaways
  • → A 50-engineer team running 8 MCP servers creates ~400 active OAuth-style credential placements. Most are long-lived, copied across machines, and invisible to security.
  • → The fix is to separate three planes that local setups conflate: inbound auth (who is calling), access control (what they can do), and outbound auth (how we reach each downstream tool).
  • → Choose OAuth grants per server, not per gateway. Authorization Code (3LO) for user-data tools like Slack and GitHub. Client Credentials (2LO) for shared backend services like an internal analytics service.
  • → Refresh tokens proactively at 80% of TTL under a per-(user, server) lock, and rotate access + refresh atomically. Reactive refresh fails the first time fifty agents wake up at 8 a.m.
  • → Token management adds ~1–4 ms p95 on a cache hit, and ~300–600 ms p95 on a cache miss requiring a provider refresh. Background refresh keeps the latter off the request path.

Un lundi matin chez Northwind. Un lundi à 8h47, l'ingénieur d'astreinte de Northwind Logistics — une plateforme de chaîne d'approvisionnement fictive mais malheureusement plausible, comptant 50 ingénieurs — reçoit un ping Slack. Leur assistant IA interne, Cargo Copilot, renvoie des erreurs 401 pour chaque appel d'outil Jira et GitHub. La cause profonde est l'ordinateur portable d'un développeur, perdu à l'aéroport pendant le week-end. Le service informatique a révoqué la session SSO de l'ordinateur portable. Mais l'ordinateur portable contenait également des jetons OAuth pour GitHub, Slack, Jira et un service d'analyse interne, chacun copié dans les paramètres de l'éditeur et les dotfiles du développeur. La sécurité ne peut pas dire avec certitude quels scopes ont été exposés, sur quelles autres machines ces jetons ont également été dupliqués (parce que les développeurs s'entraident), ni comment les révoquer tous sans perturber la matinée de tout le monde.

L'incident de Northwind n'est pas un cas isolé. C'est à cela que ressemble la gestion des justificatifs d'identité dans la plupart des entreprises six mois après une adoption enthousiaste du MCP. Le reste de cet article est l'histoire architecturale de la façon dont Northwind — et d'autres clients de TrueFoundry — cessent de reproduire cet incident.

1. Le problème de la prolifération des justificatifs d'identité : À quoi ressemblent réellement 50 développeurs × 8 serveurs MCP

Northwind compte environ cinquante ingénieurs utilisant activement des IDE compatibles MCP avec huit serveurs — GitHub, Slack, Google Workspace, Jira, un service d'analyse interne, une API logistique interne, Zendesk et un service propriétaire de planification d'itinéraires. Cela représente quatre cents connexions actives selon un calcul simple, et en pratique, un nombre considérablement plus élevé de placements de justificatifs d'identité si l'on compte les entrées de trousseau, les paramètres d'éditeur, les fichiers d'environnement shell, les variables CI et les copies que les gens font lors de l'intégration d'un coéquipier.

Le rayon d'impact d'un poste de travail compromis n'est pas d'un seul jeton. C'est un jeton multiplié par les scopes attachés à ce fournisseur, la durée de vie du jeton (souvent 90 jours ou plus pour les jetons d'accès personnels), le nombre de machines qui l'ont dupliqué, et l'absence de révocation centrale. Si un jeton OAuth Slack permet la lecture de l'historique des canaux et qu'un jeton GitHub co-localisé permet la lecture des dépôts, une seule exfiltration devient un incident inter-systèmes avant que l'équipe de sécurité n'ait eu le temps de réagir. Les justificatifs d'identité MCP locaux sont également opérationnellement invisibles : pas de piste d'audit unifiée, pas de rotation forcée, pas de vérification de politique avant l'exécution d'un outil, pas de moyen rapide de savoir qui peut appeler quoi.

Le modèle de passerelle comble cette lacune. Les développeurs s'authentifient une seule fois auprès de la passerelle, et celle-ci gère l'authentification en aval selon le modèle configuré de chaque serveur — ce qui correspond exactement à la conception décrite dans les Présentation de la passerelle MCP de TrueFoundry et les documents d'authentification et de sécurité. L'avantage n'est pas seulement la consolidation. C'est que l'authentification devient une propriété de l'infrastructure plutôt qu'une propriété de chaque ordinateur portable de l'entreprise.

2. L'architecture d'authentification à trois plans : Entrant, Contrôle d'accès et Sortant

La première décision de conception est de séparer trois plans que les configurations de développeurs locaux confondent couramment. Les documents d'authentification et de sécurité de TrueFoundry décrivent cela comme un système en trois parties : authentification entrante, contrôle d'accès, et authentification sortante. La documentation stipule explicitement que les couches entrantes et sortantes sont indépendantes, ce qui constitue un avantage architectural majeur.

  • Authentification entrante. Qui appelle la passerelle ? TrueFoundry prend en charge quatre méthodes : Clé API TrueFoundry (PAT), Jeton de compte virtuel, Jeton de fournisseur d'identité (un JWT Okta/Auth0/Azure AD validé via un fournisseur d'identitéconfiguré), et OAuth TrueFoundry (utilisé par des IDE comme Cursor et VS Code).
  • Contrôle d'accès. Que peut faire l'identité résolue ? Chaque serveur MCP dispose de Collaborateurs — utilisateurs, équipes ou comptes virtuels — avec des permissions basées sur les rôles. La portée au niveau de l'outil est obtenue en combinant des outils de plusieurs serveurs MCP en un Serveur MCP Virtuel, qui n'expose qu'un sous-ensemble sélectionné.
  • Authentification sortante. Comment la passerelle atteint-elle chaque serveur en aval ? Sept méthodes sont prises en charge : OAuth2 (Code d'autorisation), OAuth2 (Identifiants client), Clé API (Partagée), Clé API (Individuelle), Aucune authentification, Transmission de jeton, et Redirection de jeton via x-tfy-mcp-headers.

‍

Figure 1. Les trois plans de la passerelle. Une seule authentification entrante devient une ou plusieurs authentifications sortantes, régies par une décision de politique au milieu. Chaque plan est configurable indépendamment.

Le principe de conception est simple : l'identifiant utilisé pour atteindre la passerelle n'est jamais l'identifiant qui atteint l'outil en aval. La rotation entrante n'affecte aucun état OAuth du fournisseur. Un serveur MCP GitHub peut utiliser OAuth par utilisateur tandis que le service d'analyse interne utilise les identifiants client — dans le même agent, dans le même flux de requête. La passerelle est à la fois le point d'application de la politique et le courtier d'identifiants.

3. 2LO vs 3LO : Choisir le bon type d'octroi OAuth pour chaque serveur MCP

Dans le jargon d'entreprise, 2LO et 3LO désignent deux types d'octroi de RFC 6749 (OAuth 2.0), mis en correspondance avec la manière dont les utilisateurs finaux participent.

OAuth à deux pattes (2LO) utilise le flux d'octroi des identifiants client . La passerelle détient un client_id et client_secret, envoie grant_type=client_credentials au point de terminaison de jeton du fournisseur, reçoit un access_tokende courte durée, et utilise ce jeton à chaque requête vers le serveur MCP. Pas de navigateur, pas de consentement utilisateur. C'est le bon choix pour les intégrations de serveur à serveur : API d'analyse internes, microservices backend, services de données internes.

OAuth à trois pattes (3LO) utilise le Code d'autorisation . Chaque utilisateur final est redirigé vers le fournisseur (GitHub, Slack, Google) pour autoriser l'accès à ses ressources spécifiques. La passerelle stocke un refresh_tokenpar utilisateur, maintient une correspondance chiffrée de user_id → {access_token, expires_at, refresh_token, scopes, provider_metadata}, et renouvelle automatiquement le jeton d'accès.

Flow OAuth grant Who authorizes Best MCP fit
2LO Client Credentials The gateway, as a confidential client Internal analytics APIs, backend microservices, shared internal data services
3LO Authorization Code Each end user, via browser consent Slack, GitHub, Google Workspace, Atlassian, CRM tools with per-user data

Comment Northwind a fait son choix

Les huit serveurs MCP de Northwind se sont répartis à peu près équitablement. GitHub, Slack, Jira, Google Workspace et Zendesk sont tous des outils par utilisateur — la valeur de « exécuter ceci en tant qu'Alice » l'emporte sur la surcharge opérationnelle d'un jeton par utilisateur. Ceux-ci utilisent le 3LO. Le service d'analyse interne, l'API Logistique et le service de planification d'itinéraires exposent tous des données à l'échelle de l'entreprise ; aucun OAuth par utilisateur n'a de sens, et une seule subvention de Client Credentials partagée est à la fois plus simple et plus facile à auditer. Ceux-ci utilisent le 2LO.

À quoi ressemble le modèle de configuration

Remarque sur cette vue de configuration

TrueFoundry enregistre les serveurs MCP via l'interface utilisateur (Paramètres → Serveurs MCP → Ajouter un serveur MCP) plutôt que via un fichier YAML géré par le client, bien que la configuration puisse également être appliquée de manière déclarative via tfy apply. Le fichier YAML ci-dessous présente le modèle de configuration par serveur — les mêmes champs que vous remplissez dans l'interface utilisateur — en utilisant le schéma vérifié de TrueFoundry (provider-account/mcp-server-group avec des intégrations).

# Northwind's backend-group MCP Server Group: two MCP servers, one per OAuth flow.
# Schema matches TrueFoundry's provider-account / integrations pattern (tfy apply).

name: backend-group
type: provider-account/mcp-server-group
collaborators:
  - subject: team:platform-engineering
    role_id: mcp-server-manager
  - subject: team:site-reliability
    role_id: mcp-server-manager
  - subject: virtualaccount:cargo-copilot-runtime
    role_id: mcp-server-manager

integrations:

  # ---- 3LO (Authorization Code): per-user OAuth on GitHub ----
  - name: github
    type: integration/mcp-server/remote
    description: GitHub MCP server using per-user OAuth Authorization Code flow.
    url: https://github-mcp.example.com/mcp
    transport: streamable-http
    authorized_subjects:
      - team:platform-engineering
      - team:site-reliability
      - virtualaccount:cargo-copilot-runtime
    auth_data:
      type: oauth2
      grant_type: authorization_code
      authorization_url: https://github.com/login/oauth/authorize
      token_url: https://github.com/login/oauth/access_token
      client_id: tfy-secret://northwind:github-oauth:client_id
      client_secret: tfy-secret://northwind:github-oauth:client_secret
      scopes:
        - repo
        - read:org
      jwt_source: access_token
      code_challenge_methods_supported:
        - S256

  # ---- 2LO (Client Credentials): shared service identity ----
  - name: internal-analytics
    type: integration/mcp-server/remote
    description: Internal analytics MCP server using OAuth2 Client Credentials.
    url: https://analytics-mcp.internal.example.com/mcp
    transport: streamable-http
    authorized_subjects:
      - team:platform-engineering
      - team:site-reliability
    auth_data:
      type: oauth2
      grant_type: client_credentials
      token_url: https://northwind.okta.com/oauth2/default/v1/token
      client_id: tfy-secret://northwind:analytics-oauth:client_id
      client_secret: tfy-secret://northwind:analytics-oauth:client_secret
      scopes:
        - analytics:read
      jwt_source: access_token

La commande pour appliquer le fichier YAML ci-dessus est :

tfy apply -f backend-group.yaml --dry-run --show-diff

Trois choses à noter. Premièrement, tfy-secret:// est le schéma de référence FQN canonique que TrueFoundry utilise pour les secrets stockés dans un groupe de secrets TrueFoundry, avec le format séparé par des deux-points tfy-secret://<tenant>:<secret-group>:<secret-key> documenté dans le guide sur les variables d'environnement et les secrets. La valeur réside dans votre gestionnaire de secrets (AWS SSM, GCP Secret Manager, HashiCorp Vault, Azure Key Vault) ; TrueFoundry résout le FQN à l'exécution, de sorte que le fichier ci-dessus est sûr dans un dépôt Git. Deuxièmement, authorized_subjects est la liste RBAC par serveur — un mélange d'équipes et de comptes virtuels — ce qui permet au même serveur MCP de servir à la fois les développeurs humains (dont les PATs se résolvent en appartenance à une équipe) et les tâches planifiées de comptes de service (dont les jetons de compte virtuel sont mappés à virtualaccount: sujets) sans dérive de configuration. Troisièmement, l'entrée 3LO définit code_challenge_methods_supported: [S256] — PKCE est recommandé pour chaque client de type "Authorization Code" aujourd'hui et obligatoire dans la nouvelle spécification MCP du 25/11/2025.

Côté client, l'agent n'énumère jamais les modes d'authentification. Il compose une URL de passerelle, présente une seule information d'identification entrante, et la passerelle injecte la bonne information d'identification sortante par serveur :

from fastmcp import Client
from fastmcp.client.transports import StreamableHttpTransport

# Real URL pattern from the TrueFoundry docs.
transport = StreamableHttpTransport(
    url="https://llm-gateway.truefoundry.com/mcp-server/backend-group/github/server",
    headers={"Authorization": f"Bearer {user_token}"},
)

async with Client(transport) as client:
    result = await client.call_tool(
        "search_issues",
        {"query": "repo:northwind/logistics-core is:open label:critical"},
    )

4. Gestion du cycle de vie des jetons : Actualisation proactive et accès concurrentiel

La correction OAuth ne se limite pas à l'obtention de jetons. Il s'agit de les utiliser en toute sécurité sous une charge concurrente. La conception naïve actualise lorsque le fournisseur renvoie 401, ou lorsque maintenant ≥ expires_at. Cette conception fonctionne en démonstration et échoue le premier lundi où elle rencontre la réalité.

Rejouons l'incident Northwind avec une actualisation naïve : le lundi à 8h47, cinquante ingénieurs se connectent, chaque travailleur découvre indépendamment le même jeton GitHub expiré, cinquante appels parallèles POST /token frappent GitHub depuis la même passerelle, et GitHub limite le débit de l'ensemble. Cargo Copilot renvoie des 401 à tout le monde jusqu'à ce que la file d'attente se vide.

Le modèle qui tient en production a deux ingrédients : actualiser de manière proactive à 80 % du TTL, et sérialiser l'actualisation par (utilisateur, serveur) avec un verrou distribué.

def get_token(user_id, server_id):
    token = cache.get(user_id, server_id)
    if now() < token.issued_at + 0.8 * token.expires_in:
        return token

    # Crossed the 80% threshold. Serialize refresh so 50 workers
    # do not all hit the provider token endpoint at once.
    with distributed_lock(user_id, server_id):
        token = cache.get(user_id, server_id)   # re-read under lock
        if now() < token.issued_at + 0.8 * token.expires_in:
            return token   # someone else already refreshed

        new = provider.refresh(token.refresh_token)
        vault.put_atomic(user_id, server_id, new)   # access + refresh together
        cache.set(user_id, server_id, new)
        return new

L'actualisation à 80 % du TTL laisse une marge pour le décalage d'horloge, les tentatives, les points de terminaison de jetons lents des fournisseurs et les requêtes déjà en cours. Le verrou par (utilisateur, serveur) empêche un effet de "thundering herd" contre le fournisseur. La lecture à double vérification à l'intérieur du verrou n'est pas cosmétique : elle permet à la première requête d'effectuer l'actualisation tandis que les autres réutilisent le jeton fraîchement écrit.

Deux subtilités sont importantes à grande échelle. Premièrement, pour les intégrations de type "Client Credentials", la clé de cache est (server_id, provider_account); pour les intégrations de type "Authorization Code", la clé doit inclure user_id. Deuxièmement, les fournisseurs qui renouvellent les jetons d'actualisation à chaque actualisation — Google et Microsoft le font, selon la configuration — exigent que l'écriture soit atomique. Si le nouvel access_token is committed before the new refresh_token, a crash leaves the user stranded with an invalidated old refresh token and a forced re-authentication.

5. The Double-Auth Flow: A Complete Sequence Diagram

With the planes separated and the refresh path defined, an end-to-end MCP request through the gateway has a precise shape. Trace one Cargo Copilot call — "summarize the open critical-severity issues in repo northwind/logistics-core" — from agent to GitHub and back:

‍

Figure 2. Single MCP tool call traversing inbound auth, RBAC, the token vault, and (when needed) the OAuth refresh path before reaching the downstream MCP server.

Three properties of this flow are worth pointing out. The agent never receives the GitHub access token; only the gateway holds it. The MCP server never sees Northwind's inbound TrueFoundry credential unless Token Passthrough is explicitly configured (more on that in §7). And the Northwind security team gets a single governed request path with identity, authorization, and audit recorded in one place — the same property that would have let them answer the Monday-morning question, "which scopes were on that laptop?", in under a minute.

The MCP authorization specification reinforces one part of this picture: bearer access tokens must be sent in the Authorization request header and never in a URI query string. That norm matters precisely because it keeps tokens out of access logs, browser history, and proxy traces.

6. Credential Vault Integration: Why Raw Secrets Never Touch Application Memory

A production gateway should not treat OAuth refresh tokens like ordinary application rows. Client secrets and refresh tokens are long-lived keys to the kingdom — leaking a Slack refresh token is, for most purposes, equivalent to leaking the Slack workspace.

The standard pattern is envelope encryption. The gateway generates a random data encryption key (DEK) for each token record, encrypts the OAuth secret with that DEK, encrypts the DEK with a master key in KMS or a secret manager, and stores the encrypted DEK alongside the ciphertext. On read, the gateway asks the vault to decrypt only the DEK, then decrypts the token in a narrowly scoped code path that avoids logging or serializing plaintext.

In stricter deployments, token exchange itself can be pushed into a sidecar or a vault transit operation, so application code never receives raw client_secret material at all. Even with vanilla envelope encryption the blast radius changes substantially: a database dump alone is insufficient (the DEKs are not stored decrypted), and a memory dump of a general request worker is much less likely to contain broad credential material (only the keys recently requested are decrypted, and only for the duration of the request).

Northwind's deployment uses AWS KMS as the master key store and the TrueFoundry secrets store for the encrypted material. The gateway never reads a provider client_secret directly; it reads a TrueFoundry secret FQN, fetches the ciphertext, and asks KMS to unwrap. The deployment-specific work that remains — KMS key rotation policy, audit retention, operational runbooks — is the responsibility of every enterprise running the gateway, and worth a dedicated review.

7. Token Passthrough vs. Token Forwarding: Edge Cases and When to Use Each

Two patterns sound similar and have very different security implications. They are worth distinguishing carefully because misusing either one undoes most of the gateway's value.

Pattern What happens When to use Security implication
Token Passthrough The gateway forwards the inbound TrueFoundry token or IdP JWT to the MCP server unchanged. The downstream server trusts the same IdP or TrueFoundry token issuer the gateway just validated. Server sees caller identity directly. Only safe with strict audience and issuer validation, ideally with RFC 8707 resource indicators.
Token Forwarding The client supplies custom headers via x-tfy-mcp-headers, which the gateway forwards to the MCP server. The MCP server has its own auth system the gateway does not manage (legacy or custom auth). Custom headers override configured outbound auth. Should be restricted, audited, and treated as an exception path.

Token Passthrough is appropriate for internal MCP servers that validate the same enterprise identity token the gateway already validated inbound. Northwind's Logistics API is one example: the gateway and the API both trust the company's Okta tenant, so passing through the Okta JWT keeps the caller's identity intact end-to-end. The risk is token audience confusion: a token minted for the gateway should not automatically be valid for every MCP server. RFC 8707 (Resource Indicators) exists precisely to bind tokens to an intended resource, and the MCP authorization spec now requires that clients include a resource parameter and that servers validate that presented tokens were issued for them.

Token Forwarding works differently. TrueFoundry's x-tfy-mcp-headers header carries a stringified JSON keyed by the remote MCP server identifier in <group>/<server> format. Here is the format verbatim from the Virtual MCP Server docs:

import json
from fastmcp import Client
from fastmcp.client.transports import StreamableHttpTransport

extra_headers = json.dumps({
    # Custom auth for the backend-group/sentry MCP server backing this
    # virtual server. Overrides whatever outbound auth was configured.
    "backend-group/sentry": {"Authorization": "Bearer your-sentry-token"},
})

transport = StreamableHttpTransport(
    url="https://llm-gateway.truefoundry.com/mcp-server/backend-group/restricted-sentry/server",
    headers={
        "Authorization": f"Bearer {tfy_token}",   # inbound auth (gateway)
        "x-tfy-mcp-headers": extra_headers,        # outbound override
    },
)
"Custom headers from x-tfy-mcp-headers always override the default authentication. Use this sparingly, as it bypasses TrueFoundry's token management."
— TrueFoundry Authentication and Security docs

Treat Token Forwarding as an emergency hatch, not a default. The gateway's value is exactly the token management it skips.

8. Latency Benchmarks: What Token Management Adds to Every Tool Call

A gateway adds work before the downstream MCP call: inbound auth validation, policy lookup, token lookup, and occasionally a refresh. In a healthy deployment the common path is a cache hit — token metadata is valid, the encrypted token reference is warm in memory.

A note on the numbers below
The latencies in this section are engineering estimates assembled from published benchmarks of each underlying component (JWT verification, vault decrypt, OAuth provider token endpoints). They are not measured TrueFoundry production telemetry. We expose the component decomposition first so you can verify the envelope against your own stack and replace any line item with measured numbers from your deployment. The estimates are intentionally conservative; even at these envelopes, gateway overhead on the common path is dominated by the downstream MCP server call and the LLM round-trip that surrounds it.

Component costs (what we build the scenarios from)

Component Typical cost Source
HS256 JWT verify ~5–10 µs Independent JWT verification benchmarks
RS256 JWT verify ~100–200 µs WorkOS "RS256 vs HS256" analysis; same independent benchmarks
In-memory token cache fetch sub-µs to a few µs Process-local data structure access
External cache (Redis) round-trip ~1–2 ms Standard Redis intra-DC latency
AWS KMS Decrypt (warm) <10 ms typical AWS KMS guidance on decrypt latency once DEKs are cached
OAuth provider /token call ~250–400 ms average Okta Concurrency Limits docs cite 250–400 ms as their own typical API response window

Scenario envelopes

Combining the component costs above, the gateway adds roughly:

Scenario p50 added latency p95 added latency What contributes
Cache hit (valid token in memory) ~1–2 ms ~3–4 ms JWT/PAT verify (µs) + RBAC lookup (µs) + in-memory cache fetch (µs). Tail set by GC pauses and JWT key fetch caching.
Cache miss, no refresh (warm vault) ~5–8 ms ~15–25 ms Adds an encrypted-token fetch and a KMS Decrypt call (<10 ms warm per AWS). Becomes a hit on the next request.
Cache miss + provider refresh ~250 ms ~400–600 ms Synchronous POST /token to the OAuth provider, which Okta itself cites at 250–400 ms average. Tail is provider geography, TLS reuse, retries, and lock wait time.

Two takeaways. First, the common path is a cache hit, and at single-digit milliseconds it is invisible against the surrounding LLM call (typically hundreds of milliseconds to seconds) and the downstream MCP server (typically tens to hundreds of milliseconds). Second, the cache-miss-plus-refresh path is rare but expensive enough to matter when it stacks up under load — which is exactly why §4 specifies proactive refresh at 80% of TTL with a per-(user, server) lock. With proactive refresh, real users almost never see the refresh latency on the request path; the gateway has already prefetched a new token.

Every team should still instrument server-timing headers, token endpoint duration, lock wait time, vault decrypt time, and downstream MCP latency separately. The goal is to prove that token management is not the dominant cost in the agent's end-to-end tool call — the envelopes above suggest it almost never is, but you should verify with your own provider mix.

The operational fix for the cache-miss case is background refresh. When the gateway sees a token cross the 80% threshold, it refreshes on the request path once, then schedules subsequent refreshes before users notice. For high-volume shared Client Credentials integrations this almost eliminates user-visible refresh latency. For per-user Authorization Code integrations, the system can prioritize active users and frequently used MCP servers.

Before vs. after: security posture at Northwind

Dimension Before: scattered local credentials After: centralized MCP gateway
Auditability Fragmented logs across IDEs, shells, and provider dashboards. No correlation by user or tool. One governed ingress path with identity, tool, server, and outcome context. Single query answers "who called what."
Revocation speed Find every copied token on every developer machine. Hours to days, often incomplete. Revoke gateway access or remove server/tool permissions centrally. Seconds.
Rotation Manual, inconsistent, often delayed past expiry. Tokens routinely live 90+ days. Central lifecycle with proactive 80% refresh and provider-specific rotation. Atomic access + refresh writes.
Blast radius A leaked workstation may expose every configured MCP server token across every scope. One inbound token plus policy enforcement; downstream tokens stay vaulted. The leaked token can be revoked centrally.
Compliance Difficult to prove least privilege, prior consent, or who accessed what when. Audits stretch to weeks. Cleaner evidence for RBAC, approvals, usage, and credential custody. Audit log is the single source of truth.

9. FAQs

Is OAuth at the MCP layer mandatory?

No. The MCP authorization specification makes authorization optional, but HTTP-based implementations that support it should conform to the spec. In practice, enterprises with non-trivial tool estates should centralize auth long before they consider it strictly mandatory.

Should every MCP server use 3LO?

No. Use Authorization Code (3LO) when the agent acts on a human user's resources — Slack, GitHub, Gmail, CRM tools. Use Client Credentials (2LO) when the downstream server authenticates the application or gateway as itself. Mixing both within one Virtual MCP Server is normal and expected.

Does a gateway eliminate provider-side permissions?

No. It complements them. Slack, GitHub, Google, and internal APIs still enforce their own scopes and resource permissions. The gateway adds a coarser, faster layer of policy on top — server-level, tool-level, team-level — and a uniform audit trail across all of them.

OAuth 2.0 or OAuth 2.1?

The TrueFoundry gateway speaks the OAuth 2.0 grants real-world providers support today: Authorization Code (with PKCE), Client Credentials, and refresh tokens. The MCP authorization specification itself targets OAuth 2.1, which is largely OAuth 2.0 plus the security best practices that are already considered mandatory in modern deployments (PKCE everywhere, no implicit grant, no password grant). For practical purposes the two are convergent; we follow OAuth 2.1 guidance where it tightens behavior.

What about the newer 2025-11-25 MCP spec?

The November 2025 revision adds mandatory PKCE for all clients, formalizes Client ID Metadata Documents (CIMD) as the preferred client registration method, and introduces step-up authorization for incremental scope consent. These tighten security without invalidating the architecture in this post. We are tracking them and rolling support in as it matures across the ecosystem.

Where does TrueFoundry fit?

The TrueFoundry MCP Gateway is a production implementation of this architecture: a centralized server registry, configurable inbound auth, access control, multiple outbound auth modes, and token lifecycle management. This post describes the design; the Authentication and Security docs are the operational reference. For the implementation patterns covered here, see also the docs on Auth Overrides and Virtual MCP Server.

Take the next step

If you are a platform, security, or AI infrastructure leader running MCP at non-trivial scale, the highest-leverage action is to inventory every MCP server in active use, map each one to the correct inbound, policy, and outbound auth model, and decide where centralization actually wins. Northwind's mapping took two afternoons; the migration itself took a single sprint. We are happy to walk through that exercise with your team.

Read the architecture: TrueFoundry MCP Gateway authentication and security. Or book an enterprise security architecture review with our team.

Further reading

All citations in this post are linked inline. The references below collect the same URLs for printability and link-rot insurance.

Remarque : Northwind Logistics est une entreprise fictive utilisée pour ancrer la conception dans un déploiement concret. Le modèle de configuration présenté au §3 est illustratif ; les serveurs MCP de production sont enregistrés via l'interface utilisateur de TrueFoundry ou appliqués via tfy apply.

‍

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

INSCRIVEZ-VOUS
Table des matières

Gouvernez, déployez et suivez l'IA dans votre propre infrastructure

Réservez un séjour de 30 minutes avec notre Expert en IA

Réservez une démo

Le moyen le plus rapide de créer, de gérer et de faire évoluer votre IA

Démo du livre
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Découvrez-en plus

Aucun article n'a été trouvé.
October 10, 2026
|
5 min de lecture

10 meilleurs outils LLmops en 2026

comparaison
October 10, 2026
|
5 min de lecture

5 leçons sur l'exploitation d'IA agentique en production - D'après la discussion au coin du feu

Aucun article n'a été trouvé.
October 10, 2026
|
5 min de lecture

Passer à zéro dans Kubernetes : une plongée approfondie dans Elasti

Ingénierie et produits
October 10, 2026
|
5 min de lecture

L'observabilité dans les flux de travail LLM : transformer les boîtes noires en boîtes en verre

Aucun article n'a été trouvé.
Aucun article n'a été trouvé.

Blogs récents

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Faites un rapide tour d'horizon des produits
Commencer la visite guidée du produit
Visite guidée du produit