Skip to content

Architecture v0.2.0 — GCP (2026-09-28)

Everything from v0.1.0 plus a real cloud path, verified on a throwaway GKE project (9 cloud runs; final: 89 passed, 0 failed, 5 by-design skips; project deleted afterwards).

Infrastructure (deploy/terraform/gcp)

Regional GKE Autopilot cluster ramen, Firestore Native (default), Artifact Registry repo ramen, global static IP ramen-console, one groups bucket ramen-<project>-groups (per-group prefix, IAM by prefix condition), console GSA ramen-console@<project> (storage, secretmanager, container.developer, logging.viewer, IAM SA admin/user, compute security + LB admin, datastore.user, projectIamAdmin) with Workload Identity to KSA ramen-system/console.

Exposure

GKE Gateway ramen (class gke-l7-global-external-managed, static IP, self-signed TLS, routes admitted from namespaces labelled ramen.io/routes=true). HTTPRoute / → console; each zone namespace gets HTTPRoute /mcp/<group>/<zone> with a prefix rewrite to /mcp → Service worker:8080 (NEG ramen-<group>-<zone>). GCPBackendPolicy raises the console backend timeout to 300 s (IAM/compute calls exceed 30 s).

Zone = namespace

ramen-<group>-<zone>: Deployments worker and worker-canary (same image, nodeSelector on the GCP zone), Service worker (NEG), KSA worker → GSA ramen-<group>-<zone>@<project> (objectViewer on the group prefix, secretAccessor on ramen-<group>-*), Secret ramen-deploy (RAMEN_ + RAMEN_SECRET_), HPA worker. Pods get RAMEN_BUCKET_URI=gs://<bucket>/<group>; the runtime syncs it (md5 diff, stale files deleted) on every load, and the node runs the initial load at start so pods become ready without a manual reload.

Console GCP adapter (ramen_console.cloud.gcp*)

  • sync_repo clone → upload to GCS.
  • deploy(canary=True) write Secret → restart canary → wait → /admin/reload → tools/list smoke → restart stable → wait; failure scales canary to 0; job log lines streamed to the job record.
  • workers() pods + /metrics (via API-server pod proxy when the console runs outside the cluster).
  • logs() Cloud Logging k8s_container filtered by namespace (and pod), downloadable.
  • rebalance() capacity scaler 0.5 (high) / 1.0 on the zone's backend service; applied:false + note while the Gateway is reconciling, retried in the background.
  • set_ip_rules() Cloud Armor policy ramen-<group> (allow list, deny rest) attached to the backend service + RAMEN_ALLOWED_CIDRS in the Secret + worker roll; attached:false + note when the backend is busy.
  • create_service_account(), refresh(), detach_group() (namespaces + GSAs on group delete).
  • Secrets backend gcp: Secret Manager ramen-<group>-<env|all>-<zone|all>-<NAME>; store keeps only a ref.
  • Per-call authorized HTTP transports (googleapiclient is not thread-safe); ADC fallback.

Node

Separate RAMEN_ADMIN_CIDRS for /admin/* so an MCP IP lock can never block the console's deploys. RAMEN_BUCKET_URI handed to the sidecar; /admin/reload returns the load result including a sync summary.

Lessons (carried into 0.3.0)

Gateway-managed backend services report "not ready" for minutes after rollouts (retry/background everything); the LB drains old pods for ~5 s after a rollout (deploy waits, clients retry); IPv6 CIDRs must be accepted.