Deployment testing
How mosaic apps are tested against real clusters and real clouds (ADR 0010: one app, many targets).
The three workflows
| workflow | runner | what it proves | trigger |
|---|---|---|---|
e2e-k8s.yml | GitHub-hosted (k3d in Docker) | full data-mesh pipeline (backfill, CDC, poll→model, poll→file) against in-cluster peers + the rendered Helm chart | every push/PR + manual |
e2e-onprem.yml | self-hosted (mosaic-onprem) | the same harness against your k3s — own-infra production path | manual |
deploy-onprem.yml | self-hosted (mosaic-onprem) | builds a selected k8s deployment spec, imports the image into li7's k3s, and installs its Helm chart | manual |
e2e-cloud.yml | GitHub-hosted | app deployed with the rendered chart against managed stores (RDS/Cloud SQL + OpenSearch/qdrant) on AWS/GCP | manual, secrets-gated |
All three use the same entry point:
scripts/e2e/harness.sh <helm-chart-dir> <app-image> [namespace] [extra helm args...]
The harness applies the peer fixtures (scripts/e2e/peers/: postgres CDC
source, mysql poll source, clickhouse sink), seeds them, installs the chart,
port-forwards the app, and asserts:
/healthresponds;GET /api/synclists all three mirrors;- backfill lands the 2 non-draft orders in ClickHouse (
whereclause honored) and the 3 snapshotorder_items; - CDC: a new row inserted into the postgres peer appears in ClickHouse;
POST /api/sync/{mirror}/onceis accepted;- poll→file: the JSONL archive contains the CRM customers;
- poll→model:
GET /api/ordercontains the derived customers.
On any failure it dumps pods, events, and the app's logs.
Running the k3d e2e locally (optional)
The CI job does it for you; to iterate locally (needs Docker + k3d):
curl -sSL https://raw.githubusercontent.com/k3d-io/k3d/main/install.sh | TAG=5.6.0 sh
k3d cluster create e2e --image docker.io/rancher/k3s:v1.30.5-k3s1
cargo run -q -p mosaic-cli -- build examples/data-sync -o /tmp/out-ds
(cd /tmp/out-ds && docker build -f deploy/e2e-k8s/Dockerfile -t mosaic/data-sync-e2e:local app/)
docker save mosaic/data-sync-e2e:local | k3d image import -c e2e -
bash scripts/e2e/harness.sh /tmp/out-ds/deploy/e2e-k8s/helm mosaic/data-sync-e2e:local mosaic-e2e --set image.pullPolicy=Never
k3d cluster delete e2e
Registering the on-prem runner
The on-prem workflow needs a GitHub Actions runner labeled
mosaic-onprem with, on the box:
docker(for the image build),kubectlconfigured against your k3s cluster (KUBECONFIGin the runner's env),helm,- the
k3sCLI (image import viak3s ctr images import), - rustup (the render step) — or a prebuilt
mosaic-cli.
Register:
curl -sSfL https://raw.githubusercontent.com/actions/runner/main/docs/scripts/add-self-hosted-runner.sh | bash -s <URL> <TOKEN> mosaic-onprem
Then dispatch e2e-onprem from the Actions tab. The harness reuses the
mosaic-e2e namespace on your cluster — peers and the app are (re)installed
there; the image is tagged mosaic/data-sync-e2e:local, so prune old ones
with k3s ctr images rm as you like.
Deploying to the on-prem cluster
Public-app infrastructure
The public-app target is the single-node k3s cluster on li7
(192.168.11.63). It is the real on-prem hosting environment for public
Mosaic applications, separate from the disposable k3d cluster used by
e2e-k8s.yml on GitHub-hosted runners.
The network path is:
<hostname>.eisler-systems.de
-> Netcup wildcard DNS / FRITZ!Box public address
-> FRITZ!Box port forwarding (80/443)
-> li7 k3s Traefik
-> Kubernetes Ingress
-> Mosaic Service
Use a public eisler-systems.de hostname for an application Ingress and set
the class to Traefik. cert-manager is installed in the cluster and the
letsencrypt-prod ClusterIssuer obtains and renews certificates through the
Traefik HTTP-01 path. The deployment workflow therefore enables TLS and sets
the host, secret name, and ClusterIssuer without storing certificates in the
repository.
The persistent eugeis/mosaic self-hosted runner on li7 is labeled
mosaic-onprem and has Docker, Rust, Helm, kubectl, and the k3s CLI. Its
kubeconfig is local to the runner at /home/ee/.kube/config; workflows do not
upload a kubeconfig secret to GitHub-hosted runners or expose the Kubernetes
API publicly.
Dispatch deploy-onprem.yml to deploy a rendered k8s spec to the local k3s
cluster. The workflow runs only on the mosaic-onprem self-hosted runner,
which has the cluster kubeconfig installed locally. It does not use a GitHub
kubeconfig secret or expose the Kubernetes API to GitHub-hosted runners.
The workflow renders the selected app, builds its generated Dockerfile, imports
the image through k3s ctr images import, and runs helm upgrade --install.
The default inputs deploy examples/orders using its helm deployment spec to
the mosaic namespace at app.eisler-systems.de. The host input overrides
the generated chart's ingress host; cert-manager and the cluster ingress
controller remain responsible for TLS.
The self-hosted runner is trusted with cluster-admin access. Keep this workflow manual and restrict repository write access accordingly: a workflow executed on this runner can change the local cluster.
Cloud e2e (AWS / GCP)
The cloud job assumes the spec's Terraform was applied once (it creates the cluster, the registry, the managed stores, the namespace, and the DSN secrets the chart reads). Per cloud:
AWS (deploy/prod-aws/)
cd examples/full-stack # after: cargo run -q -p mosaic-cli -- build examples/full-stack -o /tmp/out-fs
cd /tmp/out-fs/deploy/prod-aws/infra
terraform init && terraform apply # EKS + ECR + RDS + OpenSearch + secrets
Repo configuration (Settings → Secrets and variables → Actions):
| name | kind | value |
|---|---|---|
E2E_AWS | variable | true |
AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY | secret | a key allowed to push to ECR and kubectl against the cluster |
AWS_E2E_REGION | variable | us-east-1 (the spec's region) |
AWS_E2E_REGISTRY | variable | the ECR repository URL (terraform output ecr_repository_uri) |
AWS_E2E_KUBECONFIG | secret | the kubeconfig for the EKS cluster (single blob; aws eks update-kubeconfig --name <cluster> renders one) |
The job builds the image, pushes it to ECR, helm installs
deploy/prod-aws/helm (db + vector mode: external — the DSN secrets come
from Terraform), waits for readiness, and smokes /health plus a write
through the managed RDS event store, then uninstalls.
GCP (deploy/prod-gcp/)
cd /tmp/out-fs/deploy/prod-gcp/infra
terraform init && terraform apply # GKE + Artifact Registry + Cloud SQL + db secret
| name | kind | value |
|---|---|---|
E2E_GCP | variable | true |
GCP_SA_JSON | secret | a service-account key with container.admin + artifactregistry.writer on the project |
GCP_E2E_PROJECT | variable | the project id |
GCP_E2E_CLUSTER | variable | the GKE cluster name |
GCP_E2E_REGION | variable | europe-west1 (the spec's region) |
Note: GCP has no first-party managed Qdrant with a stable Terraform resource,
so prod-gcp runs qdrant in-cluster on GKE (managed: false); only the
postgres store is managed (Cloud SQL). The plan layer rejects managed
combinations without a stable URL resource (deploy-managed-unavailable) —
see the matrix in ADR 0010.
Azure / Alibaba
azure-aks and alibaba-ack generate the same infra/ + helm/ +
Dockerfile triplet (AKS module + ACR; ACK module + ApsaraDB RDS / Redis /
ClickHouse). There is no automated cloud job for them yet — run the AWS job
manually as a template: apply the terraform, build/push the image from
deploy/<spec>/Dockerfile, helm install deploy/<spec>/helm, smoke,
uninstall.
Cost notes
e2e-k8s(k3d) runs on a shared runner: a few minutes of CPU, no bill.- The on-prem run uses your existing cluster.
- Cloud jobs are manual only and only touch resources you provisioned;
the job uninstalls the Helm release but never destroys Terraform state.
terraform destroywhen you are done.
Troubleshooting
- App pod stuck
ContainerCreatingon a cloud target: the DSN secret is missing —kubectl -n <ns> get secret <app>-db <app>-vector. They are created by Terraform, not the chart. - ImagePullBackOff on k3d: the image import step failed or the tag
drifted — the caller passes
--set image.pullPolicy=Neverprecisely so a preloaded image wins. - CDC never lands: the peer postgres must run
wal_level=logical(the fixture sets it); check the app logs for replication errors.