Skip to content

Ramen interface contracts (v0.3.0)

Binding for all components. Change only via a PR that updates this file and every implementer.

1. Group repo contract (what users write)

See ramen-demo-mcp-group/README.md. mcp/{tools,resources,prompts}/<name>/<name>.{py,json}, utils/ importable, mcp/requirements.txt. Proto JSON: {type, name, description, callable, input:{param:{type|enum, description?}}, output:{type}, error:{type}}. Resources add uri, mime_type. Prompts have no callable; they have skill (SKILL.md) and settings (settings.json); input params become prompt arguments and {{param}} in SKILL.md is substituted. Validation errors (folder≠name, type mismatch, bad type, missing callable) are reported per package; valid packages still load. Secrets: any string argument containing {{$<group>.<VAR>}} is substituted by the runtime before the call; code may call ramen_runtime.secrets.resolve(text). Values come from env vars RAMEN_SECRET_<GROUP>__<VAR> (upper-cased) injected by the deployer. Never logged.

2. Sidecar protocol: node-rs ↔ runtime-py

Transport: newline-delimited JSON-RPC 2.0 over the child's stdin/stdout. Node spawns python -m ramen_runtime --bucket <dir>. Stderr = runtime logs (JSON lines). Methods (params → result): - runtime.load {bucket} → {tools:[{name,description,inputSchema}], resources:[{uri,name,description,mimeType}], prompts:[{name,description,arguments:[{name,description,required}]}], errors:[{package,reason}]} - runtime.call_tool {name, arguments} → {content:[{type:"text",text}], isError:bool} - runtime.read_resource {uri} → {contents:[{uri,mimeType,text}]} - runtime.get_prompt {name, arguments} → {messages:[{role:"user",content:{type:"text",text}}]} - runtime.ping {} → {ok:true} ; runtime.shutdown {} → {ok:true} Node kills the sidecar after RAMEN_SIDECAR_IDLE_SECS (default 300) idle and respawns on demand. inputSchema is JSON Schema derived from proto input (number→number, string, integer, boolean, enum→enum of strings; all params required).

3. Node surface (node-rs) — v0.3.1: gRPC, see §11 (the HTTP endpoints below were removed)

Port RAMEN_NODE_PORT (default 8080). Since v0.3.1 the port serves gRPC (§11: ramen.v1.Mcp/Call, ramen.v1.Admin/{Reload,Metrics}, grpc.health.v1.Health); the method list, metrics fields, reload semantics, auth/CIDR/config rules below still apply, carried in gRPC metadata instead of HTTP headers. - POST /mcp → §11 Mcp/Call (one JSON-RPC 2.0 message per call). Methods: initialize, notifications/initialized, ping, tools/list, tools/call, resources/list, resources/read, prompts/list, prompts/get. JSON responses only (no SSE) in v0.1.0. Protocol version 2025-06-18. - GET /healthz / GET /readyz → §11 grpc.health.v1.Health/Check (""/ramen.v1.Mcp SERVING once runtime.load succeeded; ramen.v1.Admin = process alive). - GET /metrics → §11 Admin/Metrics, JSON {inflight, total, errors, load: "low|even|high", sidecar_alive, loaded_at, packages:{tools,resources,prompts,errors}}. load = low <30% of RAMEN_MAX_INFLIGHT, high >80%. - POST /admin/reload → §11 Admin/Reload (metadata x-ramen-admin-key) → re-run pip install for requirements.txt + runtime.load. Returns the load result. Auth: metadata authorization: Bearer <key> (constant-time compare); keys in RAMEN_MCP_KEYS (comma list) reloaded on Admin/Reload; unauth → UNAUTHENTICATED (§11). RAMEN_ALLOWED_CIDRS (comma list, default 0.0.0.0/0) → PERMISSION_DENIED outside, applies to Mcp/* only. Admin/* is gated by RAMEN_ADMIN_KEY plus the separate RAMEN_ADMIN_CIDRS (default any) so an MCP IP lock can never block the console's deploys. RAMEN_VERBOSE=1 logs full request/response bodies; otherwise one JSON line per call: ts, ip, group, method, name, status, ms, key_id. Group name from RAMEN_GROUP, zone from RAMEN_ZONE, environment from RAMEN_ENV. Config precedence: RAMEN_CONFIG file (flat KEY=value / key: value) < env < <bucket>/.ramen/env-<zone> (or .ramen/env) written by the console on deploy. The deploy file may only set deploy-scoped keys (MCP keys are unioned, group/env/zone labels, verbose, timeouts); it can never change bucket, port, python path, admin key, or proxy trust. Extra env: RAMEN_LOG_FILE (log mirror the console Logs page tails), RAMEN_CALL_TIMEOUT_SECS. No RAMEN_MCP_KEYS and no deploy file = deny all.

4. Console (console/)

Port RAMEN_CONSOLE_PORT (default 8000; TLS terminated by ingress/LB; local compose serves https on 8443 with a self-signed cert). Storage interface ramen_console.storage.base.Store (async): get/put/delete/list(collection, filters) + transaction. Collections: users, groups, environments, zones, workers, secrets, api_keys, audit, activity, config, backups. Adapters: memory (tests), firestore (honors FIRESTORE_EMULATOR_HOST), dynamodb (tests use moto). Selected by RAMEN_STORE=memory|firestore|dynamodb. Fields marked sensitive (password hashes, secret values, API key hashes, github tokens) are Fernet-encrypted with RAMEN_FERNET_KEY before write. Cloud interface ramen_console.cloud.base.Cloud: sync_repo(group, repo_url, ref, token) -> bucket_uri, deploy(group, env, zone, canary=True), rebalance(group, zone), workers(group, zone) -> [{id, load, metrics}], logs(group, zone, worker=None, tail=500), set_ip_rules(group, zone, cidrs), create_service_account(group, zone), refresh(). Adapters: local (filesystem bucket at RAMEN_BUCKET_ROOT/<group>, deploy = write env file + Admin/Reload on each worker via ramen_console.grpcclient (§11), logs = worker container stdout file), gcp (§7), aws (§8, untested on a real account). Selected by RAMEN_CLOUD=local|gcp|aws. Roles: super_admin, group_admin (per group), viewer (per group). Bootstrap super admin from RAMEN_ADMIN_EMAIL/RAMEN_ADMIN_PASSWORD on first start. Sessions: signed cookie. Passwords: argon2. API keys: rmn_<id>_<secret>, stored hashed, scoped to role+groups, header X-Ramen-Api-Key on /api/v1/*. Pages (left sidebar): Dashboard (zone→group load map, blue/green/red), Groups, Environments, Zones/Workers, Secrets, Users, API Keys, Logs, Audit, Backups, Config. Theme: coral #F26B3A on #F4F1EC, logo at /static/logo.png. HTMX for partial refresh; no JS build step. Every mutating request writes audit {ts, user, ip, action, target, ok, tags}. Probes: GET /healthz 200 always; GET /readyz 200 once the store answers and a super admin exists, else 503. Additive details settled in v0.1.0: Cloud.deploy(..., config: dict) carries the RAMEN_ / RAMEN_SECRET_ vars; local adapter writes <bucket>/.ramen/env and .ramen/env-<zone>; local workers are addressed via RAMEN_LOCAL_WORKERS="group/zone=host:port|...,..." with RAMEN_WORKER_URL fallback (default worker:8080; a scheme is accepted and stripped, https:// = TLS); permission requests live in activity.

4a. Console routes (binding; mirrored by tests/src/ramen_tests/console.py)

key route notes
login/logout POST /login form {email,password} → 303 + ramen_session cookie; /logout 401 on bad password
me GET /api/v1/me
users /api/v1/users, /api/v1/users/{id} POST {email,password,role,groups} → 201
groups /api/v1/groups, /api/v1/groups/{group} POST {name,repo_url,ref} → 201
zones /api/v1/zones, /api/v1/zones/{zone} (GET/DELETE) POST {name,provider,region}, super admin
environments /api/v1/groups/{group}/environments[/{env}] (GET/PUT/DELETE on the item), /api/v1/environments?group= POST
deploy POST /api/v1/groups/{group}/environments/{env}/deploy {canary,zone?} → 202 {id,status} poll GET /api/v1/jobs/{id}
workers GET /api/v1/groups/{group}/zones/{zone}/workers → {live:[{load}],count,size} scale count via PUT (size super admin only)
rebalance / ip-rules POST .../zones/{zone}/rebalance, PUT .../zones/{zone}/ip-rules
secrets /api/v1/groups/{group}/secrets[/{id}] POST {name,value,env?,zone?} → 201 {id,name}; value never returned
mcp-keys /api/v1/groups/{group}/mcp-keys[/{id}] POST {name} → 201 {id,key:"rmk_..."}; written to workers on deploy
api-keys /api/v1/api-keys[/{id}] POST {name,role?,groups?} → 201 {id,key:"rmn_..."}
audit / logs / dashboard GET /api/v1/audit, /api/v1/logs, /api/v1/dashboard
backups GET /api/v1/backups; POST /api/v1/backups {target} → 201 {id,release_version}; GET /api/v1/backups/{id}/download super admin
config / refresh /api/v1/config (GET), POST /api/v1/config/reload, GET|PUT /api/v1/config/sa-rules, POST /api/v1/refresh super admin
service account POST /api/v1/groups/{group}/zones/{zone}/service-account → super admin
sa-restrictions PUT /api/v1/groups/{group}/sa-restrictions 409 on clash with super-admin rules
verbose POST /api/v1/groups/{group}/environments/{env}/verbose
requests POST|GET /api/v1/requests, POST /api/v1/requests/{id}/approve permission requests (F4.2)
restore / password POST /api/v1/backups/{id}/restore, POST /api/v1/users/{id}/password
logs GET /api/v1/logs?group&zone&worker?&tail&download=1 → text/plain (+Content-Disposition attachment)
requests (v0.3.0) POST /api/v1/requests {group,zone,permission} → 201 {id,type:"permission",status}; approve → {status:"approved",applied:{ok,permissions,…}} group admin of the group; 422 unknown permission, 409 denied by super-admin/group rules (audited)
policy GET /api/v1/policy/permissions → [{permission,desc,gcp:[roles],aws:[actions]}] catalogue from policy/permissions.py
blocked PUT /api/v1/groups/{group}/environments/{env}/blocked {blocked:[names]} → env doc admin; applied by the next deploy as RAMEN_BLOCKED
config/auth GET|PUT /api/v1/config/auth {password_login?,magic_link?} → {password_login,magic_link,break_glass,providers,mail} super admin; 422 when disabling passwords would lock everyone out
auth forms GET|POST /auth/reset {email} (200 always), GET|POST /auth/reset/{token} {password} → 303 /login; POST /auth/magic {email}, GET /auth/magic/{token} → 303 + session; GET /auth/{name}/login → IdP, GET /auth/{name}/callback tokens single use (reset 24 h, magic 15 min); /auth/oauth/{name}/… kept as aliases
csrf ramen_csrf cookie (set on login and HTML pages); cookie-authenticated /api/* mutations need X-Ramen-CSRF, HTML forms the hidden csrf_token field 403 csrf token missing or invalid; X-Ramen-Api-Key requests exempt
Response shapes on gcp: rebalance {ok, load, capacity_scaler, backend_service|null, applied, note?, scaled_to?}; ip-rules {ok, policy, cidrs, backend_service|null, attached, note?}; workers live[] items carry track (stable canary), ip, phase.

5. Local stack (deploy/local/docker-compose.yml)

Services: firestore (emulator, 8081), console (8443), worker (node-rs + runtime-py in one image, 8080 gRPC h2c, §11), shared volume buckets. make demo = up, wait ready, create group demo pointing at the demo repo, deploy, call demo_calculator_tool via an MCP client (the stdio bridge, §11), print result.

6. Versioning

ramen/VERSION is the single source; Cargo.toml and both pyproject versions must equal it (CI checks).

7. GCP deployment (v0.2.0, binding)

Decisions: D6 one zone first but multi-zone capable, D16 throwaway project, D17 static IP + self-signed cert. - Terraform (deploy/terraform/gcp) creates: regional GKE Autopilot cluster ramen (var region, default us-central1), Firestore Native DB (default), Artifact Registry docker repo ramen, global static IP ramen-console, master GSA ramen-console@<project> with roles storage.admin, secretmanager.admin, container.developer, logging.viewer, iam.serviceAccountAdmin, iam.serviceAccountUser, compute.securityAdmin (Cloud Armor), compute.loadBalancerAdmin, datastore.user (Firestore), resourcemanager.projectIamAdmin (bind conditioned roles to per-group GSAs), and a Workload Identity binding (depends on the cluster so the WI pool exists) to KSA ramen-system/console. Outputs: cluster_name, region, console_ip, artifact_repo, console_gsa. Nothing per-group is created by Terraform. - Images: <region>-docker.pkg.dev/<project>/ramen/console:<VERSION> and .../ramen/worker:<VERSION>, linux/amd64, pushed by make push PROJECT=… REGION=…. - Helm (deploy/helm/ramen): namespace ramen-system holds the console Deployment (KSA console, WI-annotated), Service (NEG), the GKE Gateway described under Worker exposure (replaces Ingress/BackendConfig), HealthCheckPolicy /healthz, and the console ClusterRole. deploy/helm/ramen-worker renders one zone namespace (also rendered in Python by the console). Values: project, region, image.*, console.env. - Zone = k8s namespace ramen-<group>-<zone> created by the console when a zone is attached to a group. It holds Deployments worker and worker-canary (same image, nodeSelector: topology.kubernetes.io/zone=<gcp-zone> from the zone record's region field), Service worker (port 8080, NEG annotation, selects both), KSA worker bound to the group+zone GSA, Secret ramen-deploy (deploy env: RAMEN_ + RAMEN_SECRET_). Worker pods get RAMEN_BUCKET_URI=gs://<bucket>/<group> and RAMEN_BUCKET=/data/bucket. - Bucket sync: on runtime.load (and thus on Admin/Reload, §11) the runtime syncs RAMEN_BUCKET_URI → RAMEN_BUCKET (google-cloud-storage, ADC/Workload Identity) before loading when the URI is set. Sync is content-hash based and deletes stale files. - Console GCP adapter (ramen_console.cloud.gcp): in-cluster kube config (fallback KUBECONFIG). sync_repo → git clone/fetch to temp then upload to gs://ramen-<project>-groups/<group>/ (bucket created by Terraform var groups_bucket, one bucket, per-group prefix; IAM per group via prefix conditions). deploy(canary=True): write Secret ramen-deploy (secret values fetched from Secret Manager), kubectl rollout restart worker-canary → wait ready → Admin/Reload on the canary pod IP:8080 → Mcp/Call tools/list smoke (Health/Check SERVING when the group has no MCP keys; §11) → then restart worker → wait ready; on any failure scale canary to 0 and return the error; job log lines are streamed to the console job record. workers() = pods of worker+worker-canary + each pod's Admin/Metrics (direct pod IP; the former RAMEN_GCP_POD_PROXY mode is gone, §11). logs() = Cloud Logging resource.type="k8s_container" AND resource.labels.namespace_name="ramen-<group>-<zone>" (+ pod_name for one worker), newest N, downloadable. rebalance(group, zone) = set the backend service capacity scaler for that zone's NEG backend (0.5 when high, 1.0 otherwise) via compute API, then trigger an HPA-friendly scale if count differs. If the Gateway has not programmed a backend yet it returns 200 with applied:false and a note, never an error; if the backend service is busy (Gateway reconciling after a rollout) the capacity change is retried in the background and the response says applied:false with a pending note. set_ip_rules = Cloud Armor policy ramen-<group> (allow listed CIDRs, deny rest) attached to the worker backend service; also patched into RAMEN_ALLOWED_CIDRS of the deploy Secret. Without a programmed backend it returns 200 with attached:false and a note. Changing IP rules rolls the zone's worker pods (env comes from the Secret at pod start) so the node enforces immediately; the Cloud Armor attach retries in the background when the backend service is busy (attached:false + note). attach_zone ensures the zone identity before the first deploy (v0.3.2: it creates the GSA, its baseline grants and the WI binding when the zone KSA carries no GSA annotation — workers otherwise start with no identity and every bucket read fails 403). create_service_account(group, zone) = the same GSA ramen-<group>-<zone>@<project> with storage.objectViewer bound on the groups bucket (conditioned to the group prefix) + secretmanager.secretAccessor bound on each existing ramen-<group>-* secret (new secrets are bound at creation by the gcp secrets backend; v0.3.1 §11: no project-level bindings, the console GSA has no projectIamAdmin; project-wide roles from approved permissions need Terraform console_project_iam=true, otherwise they are reported as skipped) + WI binding to KSA ramen-<group>-<zone>/worker; extra roles only via super-admin rules (F4.2). refresh() = list namespaces/deployments/GSAs and reconcile the store. detach_group(group) (called by group delete, F4.1) deletes every namespace labelled ramen.io/group=<group> and every GSA ramen-<group>-*; failures abort the delete. A failed deploy always scales worker-canary to 0, including a canary left by an earlier deploy. - Secrets backend: RAMEN_SECRETS_BACKEND=store|gcp. gcp stores values in Secret Manager as ramen-<group>-<env>-<zone>-<NAME> (labels group/env/zone) and the store keeps only the name + resource ref. Viewers/admins never read values back through the console. - Console runtime env on GCP: RAMEN_STORE=firestore, RAMEN_CLOUD=gcp, RAMEN_SECRETS_BACKEND=gcp, RAMEN_GCP_PROJECT, RAMEN_GCP_REGION, RAMEN_GROUPS_BUCKET, RAMEN_IMAGE_WORKER. - Worker exposure (added during 0.2.0; v0.3.1 routing per §11): MCP clients reach workers through the same global LB at https://<console_ip> with gRPC metadata ramen-group/ramen-zone. GKE Gateway API: Gateway ramen in ramen-system (class gke-l7-global-external-managed, static IP, self-signed cert, allowedRoutes namespaces selector ramen.io/routes=true); console HTTPRoute for /; each worker namespace (labelled ramen.io/routes=true) gets HTTPRoute worker matching headers ramen-group=<group>, ramen-zone=<zone> + one PathPrefix rule per service in route.paths (default /ramen.v1.Mcp, /grpc.health.v1.Health and both reflection services; /ramen.v1.Admin stays cluster-internal; no rewrite) — routing Health and reflection matters because anything unmatched falls through to the console route → Service worker:8080 (appProtocol: kubernetes.io/h2c), plus HealthCheckPolicy worker of type GRPC. The LB backend service is auto-named by GKE; the console discovers it by NEG name ramen-<group>-<zone> for rebalance and Cloud Armor attachment (fallback ramen-<group>). - Console RBAC (v0.3.1, §11): KSA ramen-system/console has ClusterRole ramen-console for namespaces, networkpolicies, httproutes, healthcheckpolicies/gcpbackendpolicies, rolebindings and bind on ClusterRole ramen-console-zone only; ramen-console-zone (serviceaccounts, secrets, services, deployments, HPAs, pods read) is unbound cluster-wide and bound by the console through RoleBinding ramen-console in every ramen-<group>-<zone> namespace it attaches (a namespaced Role would be refused by RBAC escalation prevention, hence the RoleBinding→ClusterRole form). A GCPBackendPolicy sets the console backend timeout to 300s because IAM/compute operations exceed the LB's 30s default. - Sizes: worker size presets s (250m/512Mi), m (500m/1Gi), l (1/2Gi); super admin sets allowed sizes per group+zone, admins set count (HPA min/max).

8. AWS deployment (v0.3.0, binding, UNTESTED — no AWS account; D18)

Mirror of §7 with AWS primitives. Everything is built and unit-tested with moto/fakes; nothing is applied to a real account and the docs must say so. - Terraform (deploy/terraform/aws): EKS cluster ramen (one small managed node group, var region default us-east-1), DynamoDB table ramen (pk/sk single-table, on-demand; matches the existing dynamodb store adapter), S3 groups bucket ramen-<account>-groups, Secrets Manager, ECR repos ramen/console + ramen/worker, console IAM role (IRSA) with S3/SecretsManager/DynamoDB/EKS-describe/IAM-create-role (scoped by path /ramen/)/WAF/ELB permissions, AWS Load Balancer Controller + Fluent Bit → CloudWatch Logs (Container Insights) installed via Helm from Terraform, a self-signed cert imported into ACM (no domain), outputs cluster_name, region, console_url, ecr_console, ecr_worker, console_role_arn. - CloudFormation (deploy/cloudformation/ramen.yaml): the same base resources (EKS, node group, DynamoDB, S3, Secrets Manager, ECR, IAM roles, OIDC provider) as one template with parameters, for teams that cannot run Terraform. Helm steps are documented, not templated. - Helm: deploy/helm/ramen gains provider: gcp|aws. On AWS: console Ingress class alb with annotations group.name: ramen, scheme: internet-facing, certificate-arn, listen-ports [{"HTTPS":443}]; each worker namespace gets an Ingress in the same ALB group with header conditions ramen-group/ramen-zone + the route.paths set (Mcp, Health, reflection; Admin internal), backend-protocol-version: GRPC target groups and a gRPC health check (success code 0) — v0.3.1 §11; RAMEN_MCP_PATH_PREFIX is gone; worker KSA annotated with the group+zone IAM role (IRSA); pods get RAMEN_BUCKET_URI=s3://<bucket>/<group>. - Runtime: ramen_runtime.bucket.sync handles s3:// (boto3) as it does gs://. - Console AWS adapter (ramen_console.cloud.aws): sync_repo → S3 prefix; zone attach = namespace + manifests + Ingress; deploy = same canary flow as GCP; workers() = pods + Admin/Metrics (§11); logs() = CloudWatch Logs Insights query on the Container Insights log group filtered by namespace/pod, downloadable; rebalance() = weighted target groups (stable/canary) via the Ingress actions annotation; set_ip_rules() = WAFv2 IPSet + web ACL associated to the ALB (allow list, default block) + RAMEN_ALLOWED_CIDRS in the deploy Secret + worker roll; create_service_account() = IAM role ramen-<group>-<zone> trusting the cluster OIDC provider for KSA ramen-<group>-<zone>/worker, policies scoped to the group's S3 prefix and secrets path; detach_group deletes namespaces + roles; refresh reconciles. Secrets backend aws: Secrets Manager names ramen/<group>/<env|all>/<zone|all>/<NAME> with tags; store keeps sm://-style ref asm://. - Console env on AWS: RAMEN_STORE=dynamodb, RAMEN_CLOUD=aws, RAMEN_SECRETS_BACKEND=aws, RAMEN_AWS_REGION, RAMEN_GROUPS_BUCKET, RAMEN_IMAGE_WORKER, RAMEN_EKS_CLUSTER, RAMEN_ALB_GROUP=ramen, optional RAMEN_AWS_PERMISSIONS_BOUNDARY (v0.3.1: worker roles are created with the ramen-worker-boundary managed policy from Terraform/CloudFormation; the console role may only create /ramen/ roles carrying it; - disables). - Apply the §7 lessons: retry/background on ELB/WAF propagation, per-call boto3 clients (thread-safe by design), throttling backoff. - Settled during the build (deviations from the wording above): the web ACL ramen defaults to allow because the console shares the ALB; each group's IP rules add a rule that blocks the group's traffic (matched on the ramen-group header since v0.3.1) unless the source is in the group's IP set. rebalance() on AWS shifts the stable↔canary target-group weights on the zone's Ingress (there is no cross-zone capacity scaler on an ALB path rule) and returns applied:false + note until the controller has created the ALB. The worker image installs the runtime with the gcp and aws extras. apply_sa_permissions on AWS writes an inline policy ramen-sa-permissions on the worker role.

9. Auth, policy, tool blocking (v0.3.0, binding)

  • OAuth/OIDC (F6.1): providers from yaml/env (RAMEN_OAUTH_<NAME>_{ISSUER,CLIENT_ID,CLIENT_SECRET,SCOPES}); login page shows a button per provider; callback /auth/<name>/callback links or creates the user by verified email; role mapping auth.oauth.<name>.role_claim + role_map (default viewer, no groups); auth.password_login: false (super admin toggle at PUT /api/v1/config/auth) disables email/password login except for the bootstrap super admin via RAMEN_ADMIN_FORCE_PASSWORD=1 (break-glass). Audited.
  • Email auth (G12): SMTP from yaml/env (RAMEN_SMTP_{HOST,PORT,USER,PASSWORD,FROM,TLS}); invite email on user create, password reset (POST /auth/reset, POST /auth/reset/{token}), optional magic-link login (auth.magic_link: true). Dev backend RAMEN_SMTP_HOST=file://<dir> writes .eml files (tests use it). Never log message bodies.
  • SA policy engine (F4.2, F6.5): super-admin rules [{effect: allow|deny, permission: <glob>}] (existing) gate what group admins may request. POST /api/v1/requests {group, zone, permission} → super admin approve → Cloud.apply_sa_permissions(group, zone, permissions) binds the mapped cloud roles (gcp: IAM roles/conditions on the GSA; aws: IAM policy on the role; local: recorded only). Mapping table ramen_console/policy/permissions.py (e.g. bucket.read → roles/storage.objectViewer / s3:GetObject). Denied-by-rule requests are rejected with 409 and audited.
  • Tool blocking (F5.6): per environment blocked: [names] (PUT /api/v1/groups/{g}/environments/{e}/blocked); deploy writes RAMEN_BLOCKED=<comma list> into the deploy config; the node removes blocked names from tools/list, resources/list, prompts/list and answers -32601 for calls to them; the UI packages page has a block/unblock toggle per item (admins).
  • Hardening (added after the 0.3.0 security audit): signing secret = RAMEN_SESSION_SECRET → RAMEN_FERNET_KEY → a random per-process secret (never a constant; set the env in production); security headers on every response (CSP self+inline, nosniff, DENY framing, referrer same-origin; HSTS when RAMEN_COOKIE_SECURE=1); per-IP rate limit on POST /login, /auth/reset*, /auth/magic* (RAMEN_LOGIN_RATE_LIMIT, default 20/min → 429); reset/magic tokens redacted from the access log; next after login must be a same-origin path; OIDC requires email_verified: true unless RAMEN_OAUTH_<NAME>_ALLOW_UNVERIFIED=1; CSV exports neutralise formula cells; log queries accept only valid namespace/pod names (422 otherwise). Worker pods: allowPrivilegeEscalation: false, all capabilities dropped, seccomp RuntimeDefault, and a per-namespace NetworkPolicy allowing ingress only from ramen-system and the load-balancer ranges (GCP 35.191.0.0/16, 130.211.0.0/22; AWS the VPC CIDR).
  • CSRF: HTML forms carry a per-session token (ramen_csrf cookie + hidden field / X-Ramen-CSRF header); the JSON API with an API key is exempt.

10. Docs site & release material (v0.3.0, binding)

  • MkDocs Material in ramen/docs/ (mkdocs.yml at repo root, .github/workflows/pages.yml deploys on push to main and on tags to GitHub Pages). Pages: landing (logo, value proposition, quickstart, badges), how-it-works, architecture (architecture/v0.1.0.md, v0.2.0.md, v0.3.0.md + current), how-tos (local quickstart, GCP verbose, AWS verbose untested, security, secrets, devops with the API/keys), wiki (concepts: group/environment/zone/worker, protos, canary, rebalance), version tracker (versions.md generated from CHANGELOG), llms.txt and JSON-LD on the landing page for discoverability.
  • README: logo, badges, 5-command local quickstart that ends with a visible PASS (fixes fresh-user U1–U6: console URL/login, minting an rmk_ key vs rmn_ API keys, idempotent make demo with a success line), screenshots of login, dashboard, group, secrets, logs, deploy job (captured from the compose stack with Playwright into docs/img/), links to how-tos and releases.
  • ramen/skills/ cloud-ops agent skills (agentskills style): deploy-gcp, deploy-aws, rotate-keys, backup-restore, scale-zone, with a validation sub-agent note.
  • reports/launch-posts-v0.3.0.md: drafts for reddit, forums, agent portals, GitHub discussions, LinkedIn with hashtags; repo topics list. Nothing is posted by agents.

11. gRPC transport (v0.3.1, binding; supersedes §3 /mcp and §5/§7/§8 where they mention HTTP endpoints)

Decision D19. Protos: proto/ramen/v1/mcp.proto, proto/ramen/v1/admin.proto (single source; Rust via tonic-build, Python via grpcio-tools into ramen_proto packages vendored in console, runtime-py bridge and tests). - Node (ramen-node) serves one h2c port RAMEN_NODE_PORT (default 8080) with services ramen.v1.Mcp, ramen.v1.Admin, grpc.health.v1.Health (SERVING once runtime.load succeeded; NOT_SERVING before). The axum HTTP server, /mcp, /healthz, /readyz, /metrics, /admin/reload are removed. Optional node TLS: RAMEN_TLS_CERT + RAMEN_TLS_KEY (PEM) switches the port to TLS (h2); otherwise the LB terminates TLS. - Security parity: Mcp/Call requires metadata authorization: Bearer <key> validated with a constant-time compare against RAMEN_MCP_KEYS (UNAUTHENTICATED otherwise; empty key set = deny all); peer address (or x-forwarded-for only with RAMEN_TRUST_PROXY=1) must match RAMEN_ALLOWED_CIDRS (PERMISSION_DENIED); Admin/* requires x-ramen-admin-key and RAMEN_ADMIN_CIDRS; blocked names (RAMEN_BLOCKED) are filtered from */list and answered with JSON-RPC -32601; message size limit 4 MiB (tonic codec → OUT_OF_RANGE); the initial runtime.load retries every RAMEN_LOAD_RETRY_SECS (default 5, 0 disables) until it succeeds, so a worker whose cloud identity is still propagating becomes ready without an admin reload; concurrency RAMEN_MAX_INFLIGHT → RESOURCE_EXHAUSTED; admin: bad key → UNAUTHENTICATED, outside admin CIDRs → PERMISSION_DENIED, reload failure → INTERNAL; health: service "" and ramen.v1.Mcp = readiness, ramen.v1.Admin always SERVING (liveness); server reflection enabled; JSON-RPC protocol errors are returned as gRPC OK with a JSON-RPC error body; denied calls are access-logged with status: denied; one JSON access-log line per call with the same fields as before plus grpc_code. Health is unauthenticated. - Routing: clients (and the bridge) send metadata ramen-group and ramen-zone. GKE Gateway: worker Service appProtocol: kubernetes.io/h2c (fallback: node TLS + HTTP2 if the Gateway rejects h2c), HTTPRoute matches headers: [ramen-group=<g>, ramen-zone=<z>] and one PathPrefix per exposed service (route.paths: Mcp, Health, both reflection services; Admin stays cluster-internal), no path rewrite; HealthCheckPolicy type GRPC. Verified live on GKE in v0.3.2 over h2c — the TLS fallback was not needed. AWS ALB: target group backend-protocol-version: GRPC, listener rule on the two headers, health check gRPC code 0. Cloud Armor / WAF unchanged. The console targets pods directly (pod IP:8080) as before. - Console: ramen_console.grpcclient (grpcio, per-call channel or cached per target, 10s deadline) replaces every HTTP call to workers: _reload_and_smoke → Admin/Reload + Mcp/Call tools/list; workers() → Admin/Metrics; readiness → Health/Check. Local adapter identical against worker:8080. - Bridge (ramen-mcp-bridge, Python, packaged in runtime-py as a console script and in the worker image): a stdio MCP server for Claude Desktop / Cursor / the mcp SDK: ramen-mcp-bridge --target <host:port> --key <rmk_…> --group <g> --zone <z> [--tls|--insecure] [--ca <pem>]. It forwards each stdio JSON-RPC message to Mcp/Call and writes the response back; notifications are forwarded and produce nothing. Documented as the way to connect standard clients. - Local stack: compose worker exposes 8080 h2c; demo.sh and the harness call gRPC (grpcio); mcp_call.py uses the bridge. deploy/local/mcp-client-config.example.json becomes a bridge stdio config. - Harness: ramen_tests.mcp_client speaks gRPC; conformance covers UNAUTHENTICATED, PERMISSION_DENIED (CIDR), blocked -32601, size limit, health states, and the bridge end to end via the mcp SDK stdio client. - Carried security mediums fixed in the same release: GCP console GSA drops resourcemanager.projectIamAdmin in favour of iam.serviceAccountAdmin + a custom role limited to setIamPolicy on ramen-* service accounts and bucket/secret-level bindings (project-level bindings replaced by resource-level IAM on the groups bucket and secrets); AWS console role: wafv2:* narrowed to the ramen web ACL/IP sets by ARN pattern and iam:PutRolePolicy limited to /ramen/ roles; console ClusterRole loses cluster-wide secrets/serviceaccounts verbs — the console creates a namespaced Role + RoleBinding for its KSA on every zone namespace it attaches (ClusterRole keeps only namespaces, networkpolicies, httproutes and read verbs); local sync_repo passes the token via http.extraheader/GIT_ASKPASS, never in the remote URL, and strips credentials from .git/config; OIDC uses PKCE (S256) and nonce; node key compares are constant-time. - Logo: images/logo_v2.png replaces v1 in console static, docs (docs/img/logo.png, favicon), README, launch drafts; theme colours re-derived from v2. - Version: 0.3.1 (user's choice). CHANGELOG marks the transport change as breaking for HTTP MCP clients (use the bridge).

12. v0.4.0 — console UI, docs, and the verification the user asked for (binding)

Source: instructions/v0.4.md, normalised as U1–U26 in facts/v0.4.md. Decisions D21–D23.

12.1 Console UI

  • Naming (U1, U10): the product name appears once, in the logo image. The sidebar wordmark, the <title> suffix and the login heading drop it. Every button, label, heading, flash and error message is sentence case with a capital first letter ("Generate key", "Create service account", "Scale workers"); no all-caps words and no abbreviations in user-visible text — "service account", never "SA" (U2).
  • Sidebar (U6): under the logo, one row holds the signed-in email and the role side by side. The role reads Super Admin, Group Admin or Viewer and carries a per-role colour (three distinct tokens on :root). Log out sits below that row in the error colour. Version stays last.
  • Group page actions (U2): every per-zone action (scale workers, IP rules, create service account, rebalance, view logs) lives in one Actions section per zone, rendered as a button row with a single shared width class, in that order.
  • Logs (U4): two panes. Left lists entries newest first, the newest highlighted and selected by default, the newest 15 rendered eagerly and the remainder inside a scrollable frame; a worker selector above it filters the list (all by default). Right shows the selected entry's body. Every row shows consumer (the key_id, resolved to the key's name when the console knows it), timestamp, the method and tool name, and outcome as success or failure. Download keeps its current behaviour.
  • Per-zone packages (U5, U22): the group page lists tools, resources and prompts grouped by zone, each with an enable/disable toggle for that zone. Storage: environments[].blocked stays the environment-wide list (§9); a new per-zone map environments[].blocked_zones = {zone: [names]} adds to it. Deploy writes the union of both into that zone's RAMEN_BLOCKED. Route: PUT /api/v1/groups/{group}/environments/{env}/zones/{zone}/blocked {blocked:[...]}. A language model connected to a zone therefore sees only that zone's enabled packages.
  • Passwords and keys (U3): generated passwords and keys are at least 12 characters and contain upper, lower, digit and special characters; ramen_console.security.generate_password() and the key minters produce them, and setting or changing a password rejects anything weaker with a 422 naming the rule. RAMEN_MIN_PASSWORD_LEN (default 12) may raise but not lower it.
  • API keys (U7, U8, U9 / D21): the page says Generate key. Groups are picked from a multi-select of the groups the caller may grant, and are displayed as names. Every key carries client_type:
  • devops — prefix rmn_, accepted by /api/v1/* exactly as today; rejected by workers.
  • agent — prefix rmk_, accepted by a worker's ramen.v1.Mcp for the groups and zones it names; rejected by /api/v1/* with 403 {"detail":"agent key cannot call the console API"}. Generating one writes it into the deploy config of the zones it covers on the next deploy, exactly as the group MCP key does now; the group page's key form is the same control with client_type fixed to agent. Keys minted before 0.4.0 default to devops if they start rmn_, agent otherwise.
  • Dashboard (U11): the load legend and cell text name the load (low, even, high, down), never the colour. Colour stays the visual signal and every cell keeps a text label, so the grid is readable without colour.
  • Sessions (V1.4, folded in for U25): each user doc carries session_epoch; sessions embed it and are rejected when it differs. It is bumped on password change, role change, group change, delete, and on any change to config/auth.

12.2 Docs and repository

  • U12 badges: one row, one size, aligned (a single badge table or a flex row with fixed height).
  • U13: README and the wiki explain the transport plainly — JSON-RPC 2.0 over gRPC, why the node is Rust, and what secures each hop (bearer key with constant-time compare, CIDR allowlist, TLS at the edge and optionally at the node, blocked-name filtering, per-zone identity). Claims are limited to what reports/ shows; the AWS path is still described as untested.
  • U14: repository description and topics set from reports/launch-posts-v0.3.1.md.
  • U15 remove content.action.edit and edit_uri. U16 keep docs/llms.txt served, remove its nav entry. U18 dark only: one palette, no toggle. U19 toc.permalink: false so headings stop rendering a leading §.
  • U17: docs/img/architecture.svg, hand-written inline SVG, dark-mode legible, showing the real path — MCP client → bridge (stdio) → gRPC through the load balancer with ramen-group/ramen-zone metadata → Rust node → Python runtime sidecar → the group's bucket — with the console alongside writing deploy secrets and reading metrics. It replaces the ASCII shape on the how-it-works page.

12.3 Verification (U20–U26)

  • kind (U20, U21 / D22): deploy/kind/ brings up a local cluster with metrics-server, the console chart and two worker namespaces (demo/a, demo/b). make kind-up, make kind-test, make kind-down. Proves: two zones serving independently; an HPA scaling a zone's workers up under generated load and back down; rebalance changing the split; per-zone package lists differing (U22).
  • GKE (U21 second half): the same suite once on a throwaway project, then deleted, reported in reports/cloud-v0.4.0.md.
  • Transport re-check (U23): state, with evidence, what each hop is and how it is protected, including the bridge. Any hop that is plaintext by default says so.
  • Independent clients (U24 / D23): Claude Code configured against the bridge, screenshotted calling the demo tool; a Cursor config published for the user to capture. Screenshots land in docs/img/clients/.
  • Security assumptions (U25): a matrix test over roles × groups × both key types, asserting every cross-group and cross-role action is refused, and that revocation takes effect immediately.
  • Independent review (U26): a reviewer that did not write the text checks the U13 claims against the code and the reports, and files reports/claims-review-v0.4.0.md.