Canary deploys¶
POST /api/v1/groups/{group}/environments/{env}/deploy {"canary": true, "zone": optional} → 202 {id}.
Poll GET /api/v1/jobs/{id}; the group page polls it every 2 s and shows refreshing….
Steps (per zone)¶
- Sync the group repo at the environment's ref into the bucket (
gs:///s3:////buckets). GitHub token from a secret namedGITHUB_TOKEN(env- or group-scoped), else the group's stored fallback token. - Write config: the zone Secret
ramen-deploy(cloud) or<bucket>/.ramen/env-<zone>(local) withRAMEN_MCP_KEYS(all mintedrmk_keys),RAMEN_SECRET_<GROUP>__<NAME>for secrets scoped to this env/zone,RAMEN_BLOCKED,RAMEN_VERBOSE, group/env/zone labels andRAMEN_ALLOWED_CIDRS. - Canary: scale
worker-canaryto 1, rollout restart, wait for the pod to be ready (RAMEN_DEPLOY_TIMEOUT_SECS, 300 s). - Reload + smoke (gRPC to the canary pod, §11):
ramen.v1.Admin/Reloadwith metadatax-ramen-admin-key(pip install ifrequirements.txtchanged,runtime.load), thentools/listthroughMcp/Callwith the first MCP key; when the group has no keys yet,grpc.health.v1.Health/Check=SERVINGis the smoke. - Stable: rollout restart
worker, wait ready. The LB drains old pods for ~5 s; clients should retry. - Record:
last_deploy {status, at, packages, error}on the environment; audit entrydeploy.
Any failure in 3–5 (a gRPC status, a JSON-RPC error, or NOT_SERVING) scales the canary to 0 and leaves worker untouched; the error and the streamed log are in
the job and on the group page. A leftover canary from an earlier failure is also scaled to 0.
{"canary": false} skips steps 3–4 (used for the local stack and emergencies).
Rollback¶
Set the environment's ref to a previous tag/commit and deploy again. Bucket sync deletes stale files.