How it works (and why)¶
The problem¶
MCP servers are easy to write and hard to run. A team needs a place to put tools that talk to internal systems, a way to ship changes without an outage, secrets that never leak into prompts or logs, an audit trail of who called what, and a network edge that only trusted ranges can reach. Ramen is that place.
The shape¶
git repo (mcp/tools, mcp/resources, mcp/prompts)
│ deploy (console)
▼
bucket gs://…/<group> | s3://…/<group> | /buckets/<group>
│ sync on load
▼
worker pod ┌───────────────┐ JSON-RPC over stdio ┌──────────────────┐
│ ramen-node │ ───────────────────▶ │ ramen_runtime │
gRPC :8080 │ (Rust, tonic) │ ◀─────────────────── │ (Python 3.14) │
──────▶ │ auth, CIDR, │ │ loads protos, │
│ health, logs │ │ runs your code │
└───────────────┘ └──────────────────┘
▲
│ ramen.v1.Mcp/Call metadata: authorization: Bearer rmk_…, ramen-group, ramen-zone
│ LB (GKE Gateway | ALB) routes on the two headers → zone's workers
│
ramen-mcp-bridge (stdio ⇄ gRPC) ◀── Claude Desktop, Cursor, the mcp SDK
grpcurl / any gRPC client ◀── agents, CI, curl-style checks
A group is a tenant: it owns one git repo, one bucket prefix, its secrets, keys and users. An environment binds
a group to a git ref and to one or more zones; a zone is a Kubernetes namespace pinned to a cloud zone with a
worker and a worker-canary Deployment. Each worker pod runs the Rust node and, on demand, the Python runtime.
Why JSON-RPC 2.0 over gRPC (v0.3.1)¶
The MCP messages did not change; the envelope did. Every JSON-RPC request or notification travels as the body
bytes of one ramen.v1.Mcp/Call (contract §11, mcp.proto).
What that buys, compared with the HTTP endpoint 0.1.0–0.3.0 exposed:
- Binary framing and HTTP/2 multiplexing: many in-flight calls per connection, no per-request handshake, and a hard 4 MiB message limit enforced by the framework rather than by hand.
- First-class health and deadlines:
grpc.health.v1.HealthreportsSERVINGonly once the runtime has loaded code, so the LB, Kubernetes and the console all read the same signal; every call carries a deadline the node honours instead of an ad-hoc timeout header. - Typed status for the transport, JSON-RPC for the protocol: a bad key is
UNAUTHENTICATED, a blocked source range isPERMISSION_DENIED, too many calls isRESOURCE_EXHAUSTED; a blocked tool is still JSON-RPC-32601in the body. Clients can tell "the edge rejected me" from "the tool said no". - Routing by metadata, not path: the load balancer matches
ramen-groupandramen-zoneheaders, so there is no path rewrite on GCP and no path alias on AWS, and the same client config works locally and in the cloud by changing only--target.
The cost is that browsers and plain MCP-over-HTTP clients cannot connect directly. ramen-mcp-bridge (a Python
console script, also in the worker image) is a stdio MCP server that forwards each message to Mcp/Call, so
Claude Desktop, Cursor and the mcp SDK see an ordinary stdio server. See the
migration note.
Why a Rust node and a Python runtime¶
- The node owns everything that must not be slowed down or broken by user code: the gRPC surface, bearer-key
auth (constant-time compare), IP allow-lists, request bounding (
RAMEN_MAX_INFLIGHT), timeouts, health, metrics and structured logs. - The runtime owns everything users write:
pip installofmcp/requirements.txt, proto validation, argument validation, secret substitution and the call itself. It is spawned when needed and killed afterRAMEN_SIDECAR_IDLE_SECSof idleness, so a crash or a leak in tool code costs one respawn, not a pod. - They talk over newline-delimited JSON-RPC on stdin/stdout (contract §2). No sockets, no ports, nothing else to secure.
Why git → bucket → worker¶
Workers never clone. The console clones (with a GITHUB_TOKEN secret if needed, passed as a git header and never
written into the remote URL), validates, and uploads to the group's bucket prefix; workers sync by content hash on
every load. This gives one auditable deploy step, a rollback path (re-deploy an older ref), and no git credentials
on worker pods.
Why canary first¶
deploy writes the environment's config (keys, secrets, blocked tools) into the zone, restarts worker-canary,
waits for Health/Check = SERVING, calls Admin/Reload, smoke-tests tools/list through Mcp/Call, and only
then rolls worker. Any failure scales the canary back to 0 and leaves the stable track untouched. See
Canary deploys.
Why two kinds of key¶
rmk_…MCP keys are minted per group and pushed to the workers of that group on deploy. MCP clients send them as gRPC metadataauthorization: Bearer rmk_…(the bridge does this for you with--key). No keys deployed = the node denies everything.rmn_…API keys are minted per console user, scoped to a role and a set of groups, and sent asX-Ramen-Api-Keyto the console's/api/v1/*(still HTTP) for automation (CI deploys, rotation, backups). They never reach a worker.
Why the console never shows a secret¶
Secrets are stored Fernet-encrypted (or in Secret Manager / Secrets Manager) and only ever leave the console as
RAMEN_SECRET_<GROUP>__<NAME> environment variables on worker pods. Tool code references them as
{{$group.NAME}}; the runtime substitutes at call time and redacts values from errors and logs.
Why Firestore / DynamoDB and not Postgres¶
The console is a single stateless-ish node in front of a managed document store, so there is nothing to back up, patch or fail over that the cloud does not already handle. Backups of console state are JSON exports tagged with the release version, restorable at any time. See decision D2.
What "multizone HA" means here¶
The load balancer (GKE Gateway / ALB) routes calls whose metadata says ramen-group: <g> and ramen-zone: <z> to
that zone's NEG / target group, with a gRPC health check on every backend. Adding a zone adds a namespace, a
service account and a route; nothing in the data model or the deploy code assumes one zone. Rebalance adjusts
backend capacity per zone from the console; IP rules become Cloud Armor / WAF policies plus a node-level CIDR
check, so a misconfigured LB still cannot expose a worker.