Skip to main content
Fleet releases promote a canary-proven Corgtex release through GitHub Actions. Operators trigger and monitor the workflow; GitHub Actions and the Ops control plane own provider credentials, target discovery, release metadata, and health proof. The release workflow must use the GitHub environment fleet-release-production. Do not wire fleet releases to the generic production environment. Fleet promotion reuses immutable GHCR images for the release SHA. It does not build Docker images. If the canonical web or worker image for sha-<gitSha> is missing, publish it with the Release Images workflow before running a real fleet promotion. The full fleet-release.yml workflow also verifies registry availability when dispatched with dry_run=true, using its existing GHCR login and package-read permission. It inspects the candidate web and worker images and, when Ops is selected, both canonical configured Ops baseline images from the protected Railway readback. The check records linux/amd64 platform-manifest digests and immutable references, not a top-level index digest. Each inspection is bounded to 30 seconds. Missing or inaccessible images stop the workflow with a redacted error; no provider deployment, registry push, or tag change occurs in dry-run mode. This does not extend fleet-release-preflight.yml or the CLI’s default planning workflow permissions. For a real full-workflow promotion, the Ops baseline must still match the registry-check fingerprint before the first staging mutation. A missing proof, changed candidate, or baseline drift requires a fresh protected registry preflight. Existing staging drift checks remain in place. This proof covers configured baseline image availability, not database recovery or acceptance of a historical serving revision; other providers retain their existing rollback contracts. Ops source updates use the inspected candidate’s canonical repository@sha256:digest reference, not its SHA tag, so a later tag republication cannot substitute another image. Later Ops releases accept and inspect those digest-pinned configured baselines; rollback evidence retains their immutable references and rejects a mismatched descriptor. There is no automatic rollback writer in this runner. An explicitly requested older release uses the same candidate inspection and digest-pinned update path; retained baseline references remain evidence for separately governed recovery. Registry checks consume the same selected-target snapshot as promotion, including in full-workflow dry runs. Broad default/all selections exclude lifecycle-ineligible Ops targets rather than inspecting their rollback images; explicit ineligible Ops selections fail closed. Initial protected-provider preflight uses the same broad-selection eligibility rule.

Required GitHub Environment Settings

Configure these values on the fleet-release-production environment. Store secret values only in GitHub environment secrets, never in docs, issues, PR bodies, logs, or committed files.

Target Inventory Checklist

Each target variable may contain one target object or an array of target objects. Keep target inventory operational and sanitized: IDs, URLs, provider IDs, and release metadata are acceptable; raw customer content, credentials, private uploads, and bearer tokens are not. Every target must include provider: "azure" or provider: "railway"; provider is never derived from its group, hostname, or URL. Workload selectors are managed-customers, selfserve, ops, and backup-app. The legacy railway-customers and azure-selfserve selectors remain read-compatible and emit deprecation information. Railway targets must include:
  • id
  • label
  • url
  • group, one of managed-customers, selfserve, ops, or backup-app
  • provider: "railway"
  • railway.projectId
  • railway.environmentId
  • railway.webServiceId
  • railway.workerServiceId
Azure targets must include:
  • id
  • label
  • url
  • an explicit workload group
  • provider: "azure"
  • azure.resourceGroup
  • azure.acrName
  • azure.webAppName
  • azure.workerAppName
Only selfserve Azure targets are mutable through the generic fleet workflow. Azure managed-customer, core, backup, and Ops targets remain non-mutable there and never inherit selfserve resource defaults. The separate exact-target workflow below is the only repository path for one already-authoritative managed Azure production target. Set deploymentStatus or provisioningStatus to RETIRED or SUSPENDED, or set releaseEligible: false, as soon as a target cannot receive releases. Default and all selections exclude those targets. A specific workload selection reports them as blockers and a real run fails before provider mutation.

Operator Checklist

  1. Confirm fleet-release-production has the required inventory and credentials for the providers on the selected targets. Azure-only releases do not require RAILWAY_API_TOKEN; Railway-only releases do not require Azure credentials.
  2. Confirm FLEET_RELEASE_STABLE_GIT_SHA points to the canary-proven stable release, not an arbitrary main commit. Do not advance it while any blocking customer-read probe, support-connector readiness check, or required recorder smoke is pending.
  3. Confirm the Release Images workflow has published both canonical GHCR images for sha-<FLEET_RELEASE_STABLE_GIT_SHA>.
  4. For Ops and backup-app Railway targets, set both GHCR_IMPORT_TOKEN and GHCR_IMPORT_USERNAME to durable package-read credentials. Other Railway targets may use the workflow token for the immediate pull. Provider preflight verifies both protected services have a settled deployment, canonical source images, masked registry authorization, and supported startup commands. Ops requires no command override; backup-app also permits the exact role-specific npm run start --workspace=@corgtex/web or npm run start --workspace=@corgtex/worker command. Pre-deploy command overrides are rejected. Backup-app uses CORGTEX_STARTUP_MODE=web: its separate database must already contain the release migrations, and release startup does not migrate or seed it.
  5. Run a dry-run:
Dry-runs use the lightweight preflight workflow by default. That path validates configuration, resolves latest-stable, and plans rings without npm ci, Prisma generation, Docker build, Azure login, or provider mutation.
  1. Confirm the dry-run dispatches fleet-release-preflight.yml and prints each target’s workload, provider, ring, criticality, resource identifiers, deprecations, and blockers before any provider mutation. Ops must be in the final ring.
  2. If the dry-run fails on missing config, fix the named GitHub environment setting and rerun.
  3. If the preflight takes longer than 30 seconds or reports unclear blockers, stop and repair the preflight before expanding release orchestration.
  4. For a real promotion, use a specific support reason and monitor the GitHub Actions run until each target proves matching gitSha, imageTag, database=up, schema=ready, passing customer-read probes, and supportConnectorReadiness.status=ready.
  5. If any target fails, stop at the failed ring unless an operator explicitly supplies a force reason.
MISSING_SUPPORT_SCOPE means the support connector credential needs an audited scope repair. It is not customer OAuth/sign-in reauthorization.

Managed Azure Single-Target Release

managed-azure-release.yml is a manual-only path for exactly one existing managed Azure production deployment. It is not a fleet selector, provider cutover, migration-completion claim, or authorization to retire Railway. Merging the workflow does not activate it against Azure. Configure the protected GitHub environment managed-azure-release-production with required reviewers and these values: The inventory input is an opaque Build Artifact UUID from the private Corgtexdotcom/corgtex-ops repository, paired with the SHA-256 of its exact bytes. The asset must be private JSON, at most 96 KB, classified INTERNAL or CLIENT_PRIVATE, and typed DOCUMENT or OTHER. The server reloads and hashes the stored bytes, evaluates P0-05, requires exactly one ACTIVE_CLIENT_PRIMARY deployment, and verifies its opaque target identity against the requested deployment UUID. Do not put inventory bytes, target JSON, credentials, or customer-private facts in workflow inputs or logs. Before any live run:
  1. Confirm the workflow is running from the current main head and both digest-pinned release images exist for the exact 40-character release SHA.
  2. Confirm the deployment is already the authoritative managed Azure production target; this workflow does not perform a provider cutover or data migration.
  3. Confirm both Container Apps use Single revision mode, each has one expected role container, and latestRevisionName equals latestReadyRevisionName.
  4. Confirm the current and incoming application versions are mixed-version compatible for the web-then-worker update order.
  5. Name one operator who owns recovery and establish a one-writer window from baseline capture through verified success or completed compensation. Pause every other Azure writer for the two apps during that window.
  6. Dispatch the workflow with the exact inventory reference, inventory hash, deployment UUID, release SHA, release version, and audited reason. Leave execute false first.
  7. Review the bounded DRY_RUN_READY result. A dry-run reads inventory, control-plane state, source images, the destination registry, and both Container Apps, but does not acquire a lease, import images, or write Azure.
  8. Obtain separate exact-target approval through the protected environment, then dispatch the same immutable inputs with execute true while the one-writer window remains active.
Execution acquires the database-clock lease, rechecks the target, stores complete rollback templates, imports digest-pinned images, and creates fresh web then worker revisions. Each Azure operation must reach a terminal result and exact readback before progress continues. The result is one of:
  • SUCCEEDED: both apps and health evidence match the incoming release; the release record is finalized and lease capability state is cleared.
  • REJECTED: a pre-forward input or configuration rejection was confirmed before any Container App mutation, such as a source-image import rejection. The lease has been aborted or cleared, rollback templates were not applied, and operators may correct the rejected input or configuration only after returning to fresh dry-run evidence.
  • ROLLED_BACK: both apps are proven on fresh rollback revisions derived from the recorded baseline; the lease is cleared.
  • RECOVERY_REQUIRED: any drift, timeout, lost fence, crash, ambiguous ARM result, unknown partial state, or failed compensation retains the fenced lease, rollback record, and bounded recovery code. Keep the one-writer window in place, assign the named recovery owner, and do not retry the workflow as a new release. Use the fenced recovery claim only after the retained lease expires and after classifying both apps from fresh Azure reads.
If a failure occurs before any Container App mutation and leaves RECOVERY_REQUIRED, use the manual Managed Azure Release Recovery workflow before any new release run. It claims the expired fenced recovery lease through the control plane, reads the stored rollback baseline, verifies both Container Apps still match that baseline, and then finalizes rollback to clear the lease. It does not import images, patch Container Apps, or mark the incoming release successful. Any baseline drift or ambiguous Azure read must stop as RECOVERY_BLOCKED for manual investigation. Repository rollback is to disable the workflow or merge a normal protected revert. Live rollback is never inferred from an error response; it is complete only after terminal Azure operations and exact baseline readback for both apps.

Gate For Next Work

The immutable-image workflow is sufficient only if a canary promotion reuses existing images, skips Docker build work inside fleet promotion, and proves matching gitSha, imageTag, database=up, schema=ready, customer-read probes, and support-scope readiness. Move to the next release PR only when this gate is not met.