How to enable a guardrail webhook

Problem: you want a policy decision to come from something mcpgw does not have in-process — a PII classifier, a prompt-injection detector, or your own service — instead of a static allow/deny/redact rule.

Solution: name an outbound webhook under the top-level guardrails: section, then point an action: guardrail rule at it. When the rule matches, mcpgw POSTs the full JSON-RPC envelope to the webhook and maps its verdict: pass → allow, mask → forward the mutated body, reject → deny.

SECURITY: the webhook receives the full JSON-RPC body (tool arguments included). This is the first mcpgw feature that ships request bodies off-box — the audit log stays metadata-only. Treat the endpoint as a trusted data recipient: run it inside your trust boundary, over HTTPS, and mind what it logs.

Recipe

guardrails:
  - name: pii-scanner                        # [a-z0-9_-]+, unique
    url: http://127.0.0.1:8585/check         # https:// required unless loopback
    timeout: 1s                              # default 1s; max 10s
    fail_open: false                         # default: endpoint down → deny

policy:
  rules:
    - id: guard-pii-on-writes
      action: guardrail
      guardrail: pii-scanner                 # must reference a defined endpoint
      when:
        tool_prefix: "fs_"

Reload with SIGHUP — the endpoint registry swaps atomically alongside the policy engine. A reload whose guardrails: section is invalid is rejected whole; the old registry stays live.

A reference webhook server

The contract: POST with Content-Type: application/json, one attempt bounded by the configured timeout. Anything other than HTTP 200 with a well-formed verdict — non-200, timeout, malformed JSON, unknown action — is an error handled by the endpoint’s fail posture.

The request mcpgw sends:

{ "direction": "client_to_server", "method": "tools/call", "tool_name": "fs_write", "principal": "oauth:sub:alice", "body": { "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {} } }

The verdicts the webhook may return (HTTP 200 only):

{ "action": "pass" }
{ "action": "reject", "reason": "pii detected" }
{ "action": "mask", "body": { "jsonrpc": "2.0", "id": 1, "result": { "content": "[MASKED]" } } }
  • reason is optional and operator-facing only — never sent to the client, never audited.
  • mask requires body: it replaces the envelope that gets forwarded.

A ~20-line Python endpoint that rejects bodies containing a US-SSN-shaped string (plain HTTP on loopback — development only):

#!/usr/bin/env python3
import json
import re
from http.server import BaseHTTPRequestHandler, HTTPServer

SSN = re.compile(r"\b\d{3}-\d{2}-\d{4}\b")

class Guardrail(BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200)   # liveness for `mcpgw doctor`'s GET probe
        self.end_headers()

    def do_POST(self):
        raw = self.rfile.read(int(self.headers["Content-Length"]))
        req = json.loads(raw)
        hit = SSN.search(json.dumps(req.get("body", "")))
        verdict = {"action": "reject", "reason": "ssn detected"} if hit else {"action": "pass"}
        out = json.dumps(verdict).encode()
        self.send_response(200)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(out)))
        self.end_headers()
        self.wfile.write(out)

    def log_message(self, *args):
        pass

HTTPServer(("127.0.0.1", 8585), Guardrail).serve_forever()

Verify offline

policy test never makes network calls. Without a stub it reports which endpoint the rule would call; --guardrail-stub folds in the outcome a verdict would produce:

$ mcpgw policy test --config=mcpgw.yaml --tool=fs_write --guardrail-stub=reject
action:  deny
rule:    guard-pii-on-writes
guardrail: pii-scanner (stub verdict: reject)

A stubbed reject exits 1 like any deny; pass and mask exit 0. With --json, the output carries "guardrail": "pii-scanner" and a "guardrail_stub": "reject" provenance field so a stub-forced outcome is never mistaken for an engine-derived one. A stubbed mask passes the body through unchanged and notes on stderr that the webhook’s mutated body is unknowable offline.

Verify live

Start the reference server and the gateway, then send a matching call:

$ python3 guardrail_server.py &
$ mcpgw --config=mcpgw.yaml &
$ curl -i http://127.0.0.1:7332/mcp -H 'Content-Type: application/json' \
    -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"fs_write","arguments":{"note":"ssn 123-45-6789"}}}'
HTTP/1.1 403 Forbidden

{"jsonrpc":"2.0","error":{"code":-32001,"message":"policy_denied"}}

The client stays generic — the webhook’s reason never reaches it. The audit line carries the provenance:

{ "ts": "2026-07-29T00:14:24.281Z", "session_id": "demo", "method": "tools/call", "tool_name": "fs_write", "action": "deny", "rule_id": "guard-pii-on-writes", "latency_ms": 6, "guardrail_name": "pii-scanner", "guardrail_verdict": "reject", "guardrail_latency_ms": 4 }

Stop the webhook and resend: with the default fail_open: false the call is still denied, now with error_kind: "guardrail_unreachable" (and no guardrail_verdict — the failure itself is the signal). With fail_open: true the call flows and the audit line shows guardrail_verdict: "error_failopen", so the bypass is visible rather than silent.

mcpgw doctor probes each endpoint with a plain GET: any HTTP status below 500 counts as reachable and the check reports TLS validity and round-trip latency. An unreachable endpoint is an error when fail_open: false, a warning when fail_open: true. The probe does not attach configured static headers, so a 401 from an auth-gated endpoint still counts as reachable.

Production notes

  • HTTPS unless loopback. Endpoint URLs must be https:// for any non-loopback host (same posture as jwks_url). A MITM on the webhook path could forge pass verdicts and disable the guardrail entirely. tls.ca_file / cert_file / key_file cover private CAs and mTLS; tls.skip_verify exists for dev and logs a loud startup warning.
  • Credentials via env. headers values expand ${VAR} at config load: Authorization: "Bearer ${GUARDRAIL_TOKEN}".
  • Size the timeout honestly. The call is synchronous: every matched request pays the round trip, and when the endpoint hangs matched requests pay the full timeout until the breaker opens (below). Default 1s, hard cap 10s. Keep when matchers narrow so only traffic that needs the check pays for it; watch guardrail_latency_ms in audit.
  • No retries. One attempt per matched request. Retrying on the hot path would double worst-case latency on a sick endpoint. (The async audit webhook sink does retry — different tradeoff.)
  • Bound what leaves the box. max_body_bytes (default 1 MiB) caps the envelope an endpoint can be sent. Oversize bodies are refused before the callout — never truncated, since a verdict on a partial envelope approves content the guardrail never saw — and take the fail posture with error_kind: guardrail_body_too_large. The refusal does not count toward the breaker. Set 0 to remove the cap and rely only on the gateway’s 16 MiB inbound limit.
  • A sustained outage trips the breaker, not every request. After 5 consecutive failures (breaker.failure_threshold) calls short-circuit for 5s (breaker.open_duration) and apply the fail posture immediately, then one probe tests recovery. The outcome is identical to an attempted-and-failed call — you keep the enforcement, you stop paying the timeout and stop hammering a sick endpoint. Watch for circuit breaker open in the logs; guardrail_latency_ms near 0 on denied requests is the audit-side tell. Set failure_threshold: 0 to disable it.
  • SSE streams are not intercepted in v1. A direction: server_to_client guardrail rule fires on buffered JSON responses only, never per SSE frame; policy lint warns (guardrail-sse-direction) so you do not believe SSE is covered. Regex redact/deny rules remain the SSE mechanism.

Replay caveat

policy replay cannot re-derive a guardrail decision offline — the verdict is not in the audit log and request bodies never are. Records whose audited decision came from a callout (guardrail_name is set) are counted as guardrail_skipped, reported on stderr, shown inline in the text summary and as guardrail_skipped in --json, and excluded from the divergence gate. Divergences on any other record still fail the gate.