Skip to content
Blog

Hardening MCP Servers: Input Validation, Rate Limits, and Authentication

A practical hardening guide for production MCP servers: schema-validated tool inputs, payload caps, per-client rate limits with 429 semantics, OAuth 2.1 enforcement, and transport-level controls grounded in the MCP specification.

Published on • October 11, 2026

AI Assistant

An internal research agent was wired to a new MCP server that exposed a search_docs tool. The tool took a free-form query string, appended it to an internal HTTP call, and returned whatever came back. Nothing validated the query length, nothing authenticated the caller beyond a shared token in an environment variable, and nothing counted requests. Within a week, a prompt-injection payload buried in a retrieved web page convinced the agent to call search_docs with a 40 MB query string in a tight loop. The upstream service fell over, and the incident report blamed “the AI.” The bug was older than the model: an MCP server that trusted its inputs, its callers, and its traffic profile.

That failure mode is the norm, not the exception. MCP tool results flow straight into model context, tool arguments are generated by a non-deterministic system, and many teams deploy their first server with the SDK defaults. The MCP specification is explicit about what production servers owe their operators: servers MUST validate all tool inputs, implement proper access controls, rate limit tool invocations, and sanitize tool outputs. This tutorial walks through exactly those layers.

In this tutorial, you will learn how to:

  • Build a threat model for an MCP server covering tool poisoning, prompt injection via tool descriptions and outputs, oversized payloads, and unauthenticated control planes
  • Validate every tools/call argument against a strict JSON Schema before execution
  • Enforce size caps on arguments, results, and nested structures
  • Implement per-client rate limiting with token buckets, quotas, and correct 429 plus Retry-After semantics
  • Enforce OAuth 2.1 bearer authentication with audience validation and resource indicators
  • Secure the Streamable HTTP transport (Origin validation, header consistency, localhost binding)
  • Emit structured audit logs for every tool invocation
  • Run a production hardening checklist before exposing a server to real agents

Key technologies: MCP specification revision 2026-07-28, JSON Schema 2020-12, OAuth 2.1 (draft-ietf-oauth-v2-1), RFC 9728 Protected Resource Metadata, RFC 8707 Resource Indicators, Streamable HTTP transport, Python 3.12, Pydantic v2.

Prerequisites

  • Python 3.12+ and a working MCP Python SDK environment (mcp package)
  • An OAuth 2.1 authorization server you can issue test tokens against (Keycloak, Auth0, or similar)
  • Familiarity with JSON Schema and HTTP status code semantics
  • A Redis instance (or any store with atomic counters) for distributed rate limiting

Threat Model for an MCP Server

Scope the model before writing controls. An MCP server faces threats on three surfaces: what it advertises, what it accepts, and what it returns.

ThreatVectorImpactPrimary control
Tool poisoningMalicious or compromised tool description steers the modelSilent data exfiltration, rogue tool callsDescription review, signed server manifests, human-in-the-loop confirmation
Prompt injection via outputsTool result contains adversarial instructionsAgent obeys attacker instead of userOutput sanitization, size caps, treating results as untrusted data
Oversized payloadsMulti-megabyte arguments or resultsMemory exhaustion, context flooding, denial of walletPer-argument and per-result byte limits
Unauthenticated controlOpen tools/call endpointFull tool surface abuseOAuth 2.1 bearer enforcement, no token passthrough
Argument injectionType confusion, path traversal, SQL fragments in argumentsBackend compromiseJSON Schema validation, allowlists, parameterized queries
State handle hijackingGuessing a cart/workflow handle returned by a toolCross-user data accessBind handles to the verified token subject server-side
Traffic abuseLoops, retry storms, credential stuffingOutage, quota burnPer-client rate limits and quotas

The specification’s security best practices page documents the same classes of attacks in depth: confused deputy via proxy servers, token passthrough, SSRF through OAuth metadata URLs, and state handle hijacking. Read it as your red-team checklist.

Input Validation: Schema First, Code Second

MCP tools declare an inputSchema. In the 2026-07-28 revision the schema defaults to JSON Schema 2020-12, MUST be a valid JSON Schema object, and tool names SHOULD stay within [A-Za-z0-9._-] and 1–128 characters. Treat that schema as a contract you enforce, not documentation you hope the model reads.

Validation happens at two boundaries:

  1. Transport boundary — reject malformed JSON-RPC before any handler runs. Unknown tools and malformed requests are protocol errors (-32602), returned as JSON-RPC errors.
  2. Tool boundary — validate arguments against the declared schema. Input validation failures are tool execution errors: return a result with isError: true and actionable text so the model can self-correct. The specification explicitly wants validation feedback in this form rather than as opaque protocol failures.
from typing import Any
from jsonschema import Draft202012Validator
from mcp.server.fastmcp import FastMCP, Context

mcp = FastMCP("docs-search")

SEARCH_SCHEMA: dict[str, Any] = {
    "type": "object",
    "properties": {
        "query": {"type": "string", "minLength": 1, "maxLength": 512},
        "limit": {"type": "integer", "minimum": 1, "maximum": 50, "default": 10},
        "collections": {
            "type": "array",
            "items": {"type": "string", "enum": ["engineering", "legal", "hr"]},
            "maxItems": 3,
        },
    },
    "required": ["query"],
    "additionalProperties": False,
}

_validator = Draft202012Validator(SEARCH_SCHEMA)

def validate_args(arguments: dict[str, Any]) -> dict[str, Any]:
    errors = sorted(_validator.iter_errors(arguments), key=lambda e: e.path)
    if errors:
        detail = "; ".join(f"{'/'.join(map(str, e.path)) or '<root>'}: {e.message}"
                           for e in errors[:5])
        raise ValueError(f"Invalid arguments: {detail}")
    return {
        "query": arguments["query"].strip(),
        "limit": arguments.get("limit", 10),
        "collections": arguments.get("collections", []),
    }

Three rules make this hold in production:

  • additionalProperties: false everywhere. Models pad arguments with plausible-looking extras; unknown keys are a signal of prompt drift or injection, not something to ignore.
  • Enums and numeric ranges, not free strings. An enum on a collection name eliminates an entire class of injection.
  • Fail as tool execution errors. Return isError: true with a message that names the field and the bound: "limit: must be <= 50". The model recovers; a bare -32602 does not.

Size Caps and Payload Hygiene

Schema validation constrains shape, not weight. A schema-valid string can still be 20 MB. Enforce byte caps explicitly at the transport layer before parsing deep structures, then again per field:

SurfaceSuggested capOn violation
Raw HTTP body256 KB–1 MBHTTP 413 before JSON parse
Single string argument8–64 KBTool execution error
Array itemsmaxItems in schemaTool execution error
Base64 blobs (images)Per-tool budgetTool execution error
Tool result (unstructured)64–256 KBTruncate + flag, or resource link
structuredContent256 KBResource link instead of inline

Truncating silently is worse than rejecting: a model that receives a truncated JSON fragment will hallucinate the rest. Prefer returning a resource_link content block so the client fetches the full payload out-of-band, keeping model context bounded.

Also sanitize outputs. Tool results are untrusted data that re-enter the model. Strip control characters, normalize newlines, and consider wrapping third-party HTML or markdown in a sanitizer before returning it. The specification’s tool security considerations require servers to sanitize outputs for exactly this reason.

Rate Limiting and Quotas

The specification requires servers to rate limit tool invocations but leaves the algorithm open. For agent traffic, a token bucket per client subject with a secondary daily quota covers both bursty exploration and runaway loops.

Key rate limits by the verified token subject (sub claim or client ID), never by IP alone — agents run behind NAT and serverless egress pools.

import time
import redis

r = redis.Redis(host="localhost", port=6379, decode_responses=True)

class RateLimiter:
    """Token bucket: capacity burst, refill_per_sec sustained rate."""

    def __init__(self, capacity: int, refill_per_sec: float):
        self.capacity = capacity
        self.refill = refill_per_sec

    def check(self, key: str) -> tuple[bool, int]:
        now = time.time()
        bucket = f"rl:{key}"
        tokens, ts = r.hmget(bucket, "tokens", "ts") or (self.capacity, now)
        tokens = float(tokens) if tokens is not None else float(self.capacity)
        ts = float(ts) if ts is not None else now
        tokens = min(self.capacity, tokens + (now - ts) * self.refill)
        if tokens < 1:
            retry_after = max(1, int((1 - tokens) / self.refill))
            return False, retry_after
        r.hset(bucket, mapping={"tokens": tokens - 1, "ts": now})
        r.expire(bucket, 3600)
        return True, 0

limiter = RateLimiter(capacity=20, refill_per_sec=2)

Return rate-limit failures as HTTP 429 with a Retry-After header (RFC 6585). Keep the semantics distinct from authorization failures:

ConditionStatusBody / header
Missing or invalid token401WWW-Authenticate: Bearer resource_metadata="..."
Valid token, missing scope403WWW-Authenticate: ... error="insufficient_scope", scope="..."
Over quota429Retry-After: <seconds>
Malformed request / header mismatch400JSON-RPC error, e.g. -32020

Two agent-specific details matter:

  • Retry storms amplify. Agents retry 429s less intelligently than humans expect. Always send Retry-After, and consider exponential server-side backoff keys per subject so a looped client’s effective rate decays.
  • Gateways can enforce without parsing bodies. The Streamable HTTP transport mirrors tools/call metadata into Mcp-Method and Mcp-Name headers, and tool schemas can annotate parameters with x-mcp-header so values surface as Mcp-Param-* headers. A WAF or load balancer can rate limit by tool name or tenant parameter before the request reaches your process — but only after verifying header/body consistency (more on that below).

Authentication Enforcement

MCP authorization is OAuth 2.1-based: the MCP server acts as a resource server, the client as an OAuth client, and the authorization server issues the tokens. Servers MUST implement OAuth 2.0 Protected Resource Metadata (RFC 9728) so clients can discover the authorization server, and clients discover it from the resource_metadata URL in a 401 WWW-Authenticate challenge.

The enforcement rules that matter on the server:

  1. Validate every request. Authorization MUST be included in every HTTP request; tokens in query strings are forbidden.
  2. Validate the audience. Servers MUST confirm the token was issued for them (RFC 8707 resource indicators). A token minted for another service is not a valid credential, even if the signature verifies.
  3. Never pass tokens through. Token passthrough — accepting any bearer token and forwarding it downstream — is explicitly forbidden. It breaks rate limiting, audit trails, and trust boundaries.
  4. Return precise challenges. Include scope in WWW-Authenticate so clients request only what the operation needs.
import jwt
from jwt import PyJWKClient

jwks_client = PyJWKClient("https://auth.example.com/.well-known/jwks.json")
RESOURCE = "https://mcp.example.com"

def authenticate(headers: dict) -> dict:
    auth = headers.get("authorization", "")
    if not auth.startswith("Bearer "):
        raise PermissionError("missing bearer token")
    token = auth.removeprefix("Bearer ")
    key = jwks_client.get_signing_key_from_jwt(token).key
    claims = jwt.decode(
        token, key, algorithms=["RS256"],
        audience=RESOURCE,          # RFC 8707 audience binding
        issuer="https://auth.example.com",
    )
    return claims

Scopes are your least-privilege dial: request only docs:search for the search tool, challenge with error="insufficient_scope" when a write tool is called without docs:write, and never publish an omnibus * scope. The security best practices page’s scope minimization section is blunt about the blast radius of broad tokens.

Transport Security

The Streamable HTTP transport (the single POST endpoint that replaced HTTP+SSE) carries its own requirements:

  • Servers MUST validate the Origin header on incoming connections to prevent DNS rebinding, responding 403 on invalid origins.
  • Local servers SHOULD bind to 127.0.0.1 only.
  • Every POST MUST carry MCP-Protocol-Version; a mismatch with the body _meta value is a 400 with JSON-RPC error -32020 (HeaderMismatch).
  • Servers MUST reject requests where Mcp-Method / Mcp-Name / Mcp-Param-* headers disagree with the body — intermediaries may route or rate limit on those headers, so a mismatch is a policy bypass attempt.

Terminate TLS at the edge, disable HTTP/1.0 keep-alive tricks that leak connections, and set explicit timeouts on both the JSON-RPC handler and any outbound call a tool makes. The specification asks clients to implement tool-call timeouts; servers should not be the component that hangs.

Logging and Audit

Log every invocation as a structured event before execution and again on completion:

import json, logging, time

audit = logging.getLogger("mcp.audit")

def audit_event(event: str, ctx: Context, tool: str, **fields):
    audit.info(json.dumps({
        "event": event,
        "tool": tool,
        "ts": time.time(),
        "request_id": getattr(ctx, "request_id", None),
        "subject": getattr(ctx, "subject", None),
        **fields,
    }))

Minimum fields: timestamp, request ID, verified subject, tool name, argument digest (hash, not raw PII), outcome (ok / error / rate_limited / denied), duration, and result size. Rate-limit denials and scope challenges belong in the same stream — they are your earliest signal of an agent loop or a credential problem.

Common Pitfalls

  • Validating only at the edge. A gateway schema check without in-process validation misses bugs in direct-to-service traffic.
  • Rate limiting by IP. Serverless agents share egress IPs; you will throttle innocent tenants and miss distributed abuse. Key by token subject.
  • Confusing 401, 403, and 429. Agents treat them differently. 401 means re-authenticate, 403 means fix scopes, 429 means back off. Collapsing them into 500 causes retry storms.
  • Trusting tool annotations. Clients MUST treat annotations from untrusted servers as untrusted; servers should treat their own annotations as display hints, never as security policy.
  • Accepting state handles as authentication. A cart ID from a tool result is a name, not a credential — re-check the caller’s authorization against it on every call.
  • Skipping output sanitization. Input validation alone does not stop injection delivered out of a tool back into the model.

Production Hardening Checklist

  • Every tool has a strict JSON Schema with additionalProperties: false, enums, and numeric bounds
  • Arguments are validated in-process; failures return isError: true with actionable messages
  • Raw body, string, array, and result size caps enforced before and after execution
  • OAuth 2.1 bearer tokens validated on every request with audience (resource) and issuer checks
  • Token passthrough disabled; downstream services receive server-minted credentials only
  • Per-subject token-bucket rate limits plus daily quotas; 429 responses include Retry-After
  • Origin header validated; local deployments bind localhost only
  • Header/body consistency (MCP-Protocol-Version, Mcp-Method, Mcp-Name) verified
  • Structured audit log covers invocations, denials, and scope challenges
  • Human-in-the-loop confirmation enabled for mutating tools

Further reading