Architecture v0.2.0 — GCP (2026-09-28)¶
Everything from v0.1.0 plus a real cloud path, verified on a throwaway GKE project (9 cloud runs; final: 89 passed, 0 failed, 5 by-design skips; project deleted afterwards).
Infrastructure (deploy/terraform/gcp)¶
Regional GKE Autopilot cluster ramen, Firestore Native (default), Artifact Registry repo ramen, global
static IP ramen-console, one groups bucket ramen-<project>-groups (per-group prefix, IAM by prefix condition),
console GSA ramen-console@<project> (storage, secretmanager, container.developer, logging.viewer, IAM SA
admin/user, compute security + LB admin, datastore.user, projectIamAdmin) with Workload Identity to KSA
ramen-system/console.
Exposure¶
GKE Gateway ramen (class gke-l7-global-external-managed, static IP, self-signed TLS, routes admitted from
namespaces labelled ramen.io/routes=true). HTTPRoute / → console; each zone namespace gets HTTPRoute
/mcp/<group>/<zone> with a prefix rewrite to /mcp → Service worker:8080 (NEG ramen-<group>-<zone>).
GCPBackendPolicy raises the console backend timeout to 300 s (IAM/compute calls exceed 30 s).
Zone = namespace¶
ramen-<group>-<zone>: Deployments worker and worker-canary (same image, nodeSelector on the GCP zone),
Service worker (NEG), KSA worker → GSA ramen-<group>-<zone>@<project> (objectViewer on the group prefix,
secretAccessor on ramen-<group>-*), Secret ramen-deploy (RAMEN_ + RAMEN_SECRET_), HPA worker.
Pods get RAMEN_BUCKET_URI=gs://<bucket>/<group>; the runtime syncs it (md5 diff, stale files deleted) on every
load, and the node runs the initial load at start so pods become ready without a manual reload.
Console GCP adapter (ramen_console.cloud.gcp*)¶
sync_repoclone → upload to GCS.deploy(canary=True)write Secret → restart canary → wait →/admin/reload→tools/listsmoke → restart stable → wait; failure scales canary to 0; job log lines streamed to the job record.workers()pods +/metrics(via API-server pod proxy when the console runs outside the cluster).logs()Cloud Loggingk8s_containerfiltered by namespace (and pod), downloadable.rebalance()capacity scaler 0.5 (high) / 1.0 on the zone's backend service;applied:false+ note while the Gateway is reconciling, retried in the background.set_ip_rules()Cloud Armor policyramen-<group>(allow list, deny rest) attached to the backend service +RAMEN_ALLOWED_CIDRSin the Secret + worker roll;attached:false+ note when the backend is busy.create_service_account(),refresh(),detach_group()(namespaces + GSAs on group delete).- Secrets backend
gcp: Secret Managerramen-<group>-<env|all>-<zone|all>-<NAME>; store keeps only a ref. - Per-call authorized HTTP transports (googleapiclient is not thread-safe); ADC fallback.
Node¶
Separate RAMEN_ADMIN_CIDRS for /admin/* so an MCP IP lock can never block the console's deploys.
RAMEN_BUCKET_URI handed to the sidecar; /admin/reload returns the load result including a sync summary.
Lessons (carried into 0.3.0)¶
Gateway-managed backend services report "not ready" for minutes after rollouts (retry/background everything); the LB drains old pods for ~5 s after a rollout (deploy waits, clients retry); IPv6 CIDRs must be accepted.