# Distributed Execution: Readiness Checklist Pre-run validation for a workflow's execution target. Companion to the [user guide](distributed-execution-guide.md); this document covers AC3. Run these checks **before the first execution** of a workflow, and again after any change to the run launcher, the executor, the target namespace or the code location image. > **L4–L10 were cleared on sandbox-cat-dat on 2026-08-31.** Checks L1–L3 and L11, > which exercise the payload contract, the message-parsing path and image-level > isolation, are verified locally. The **tightly coupled** Kubernetes path has > been assessed against the same cluster but not executed — T2, T3, T5 and T8 > are confirmed from the live instance configuration and C6 fails there. See > [section 2.1](#21-sandbox-state-2026-08-31). --- ## 1. Common checks — both setups | # | Check | How to verify | Expected evidence | |---|---|---|---| | C1 | Code location loads | Dagster UI → **Deployment** → **Code locations** | Location `distributed-execution` shows status *Loaded*, with a recent load timestamp and no error banner | | C2 | Jobs are registered | Dagster UI → **Jobs** | The jobs listed in the guide's section 5.5 appear under the code location | | C3 | Image tag matches the intended release | `kubectl -n dagster get deploy -l dagster/code-location=distributed-execution -o jsonpath='{.items[*].spec.template.spec.containers[*].image}'` | Tag equals the version in `pipeline.variables.sh`; never `latest` | | C4 | Image architecture matches the nodes | `docker manifest inspect ` | Includes `linux/amd64`; a manifest with only `linux/arm64` produces `no match for platform` at pull time | | C5 | Run launcher type is as intended | `kubectl -n get cm dagster-instance -o yaml` | `run_launcher` block shows `K8sRunLauncher` and the expected `job_namespace` | | C6 | Target namespace exists and is schedulable | `kubectl get ns ` and `kubectl -n get resourcequota` | Namespace is `Active`; remaining quota exceeds the job's aggregate requests. Also confirm the launcher's `instance_config_map`, `postgres_password_secret` and any PVC volumes exist **in that same namespace** — these references do not cross namespaces, and on sandbox-cat-dat they do not match `job_namespace` (see section 2.1) | ## 2. Tightly coupled checks | # | Check | How to verify | Expected evidence | |---|---|---|---| | T1 | Run pod reaches the metadata database | Launch `tightly_coupled_in_process_job` | Run reaches `SUCCESS`; `report_execution_target` output metadata lists `DAGSTER_POSTGRES_HOST`, `DAGSTER_POSTGRES_USER` and `DAGSTER_POSTGRES_DB` under `env_vars_present`, and `env_vars_missing` is empty | | T2 | Vault injection works | Same run; inspect the run pod | `kubectl -n dagster describe pod ` shows the `vault-env` init container completed; no `vault:` literal remains in the process environment | | T3 | Object storage is reachable | Same run, if the workflow uses S3 | No `EndpointConnectionError` in run logs; `S3_ENDPOINT_URL` present in `env_vars_present` | | T4 | Multiprocess fan-out actually fans out | Launch `tightly_coupled_local_job` | Run succeeds; `summarise_results` metadata shows one entry per unit in `contributing_workers` and a single entry in `contributing_hosts` — separate processes, same machine | | T5 | RBAC permits step Jobs | `kubectl -n dagster auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev` | Returns `yes`; required only for `k8s_job_executor` | | T6 | Step pods are actually created | Launch `tightly_coupled_k8s_job`, then `kubectl -n dagster get jobs -l dagster/run-id=` | One Job per mapped unit; `contributing_hosts` now shows one entry **per unit**, not one | | T7 | Step pod egress is permitted | Same run | Steps do not hang in `STARTING`; run logs contain no connection timeouts to port 5432 | | T8 | Failure surfaces as a pod failure | Force a step failure in a scratch namespace | `failPodOnRunFailure: true` is set, and the step pod reports `Failed` rather than `Completed` | ### 2.1 Sandbox state, 2026-08-31 The T checks were assessed against sandbox-cat-dat by reading the live `dagster-instance` ConfigMap in `dataprovider01`. None of them could be *executed*: the tightly coupled Kubernetes path needs this code location registered on the platform Dagster, which is blocked on the control plane skew, and unlike the pipes transport it cannot be exercised from a standalone pod — every step pod connects to the metadata database on its own, so they must share instance storage rather than a local SQLite file. What the configuration shows: | # | Finding | |---|---| | C5 | `run_launcher` is `K8sRunLauncher` with `job_namespace: dagster` | | T2 | Vault injection is configured — `pod_template_spec_metadata` carries the banzaicloud annotations with role `sandbox-cat-dat-role` | | T3 | `S3_ENDPOINT_URL` is `https://s3.sandbox-cat-dat.simpl-europe.eu`, with access keys injected from Vault | | T5 | Passes. `dagster-role`, bound to `dagster-svc-account`, grants `batch/jobs` with create, delete, get, list, patch, update and watch | | T8 | `fail_pod_on_run_failure: true` is set | **C6 fails, and it is the reason nothing has ever run here.** The launcher sends run pods to namespace `dagster`, but every namespace-local dependency it names — `instance_config_map: dagster-instance`, `postgres_password_secret: dagster-postgresql-secret`, and the `dagster-shared-pvc` volume — exists in `dataprovider01`. A ConfigMap, Secret or PVC reference resolves only within the pod's own namespace, so a run pod placed in `dagster` would reference three objects that are not there. Consistent with that, `dataprovider01` holds no pods labelled `dagster/run-id` and the daemon log shows no launch activity. One caveat on the evidence: project-scoped access means `kubectl get ns dagster` returns `Forbidden` rather than `NotFound`, so whether that namespace exists is unconfirmed. It does not change the conclusion — the dependencies are in the wrong namespace either way. This is a platform configuration defect, not a defect in this service. It is exactly what C6 exists to catch. ## 3. Loosely coupled checks | # | Check | How to verify | Expected evidence | |---|---|---|---| | L1 | Payload contract is intact | `uv run pytest tests/test_loosely_coupled.py` | `test_payload_does_not_import_dagster` passes — the payload imports `dagster_pipes` only | | L2 | Pipes round trip works | Launch `loosely_coupled_subprocess_job` | Run reaches `SUCCESS`; run logs contain the payload's `External payload started on …` line, proving `pipes.log` crossed the channel | | L3 | Silence is treated as failure | Same test module | `test_silent_message_path_is_treated_as_failure` passes — an empty message list raises rather than yielding an empty result | | L4 | Payload image is pullable by the target cluster | `kubectl -n run pull-probe --image= --restart=Never --command -- true` | Pod reaches `Completed`; no `ImagePullBackOff`. Both Gitea images pull anonymously — a bare registry `GET` returns 401, but that is the start of the Docker token handshake, not a refusal | | L5 | Dispatcher can create Jobs in the payload namespace | `kubectl -n auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev` | Returns `yes` | | L6 | Dispatcher can read pod logs — **the message channel** | `kubectl -n auth can-i get pods/log --as=system:serviceaccount:dagster:dagster-dev` | Returns `yes`. A `no` here breaks reporting *without* failing the workload | | L7 | Payload Job is actually created | Launch `loosely_coupled_k8s_job`, then `kubectl -n get jobs -l app.kubernetes.io/name=distributed-execution-payload` | One Job per dispatch, labelled `dagster/execution-target=loosely-coupled` | | L8 | Payload has **no** orchestration connectivity | Same run; inspect run logs | No `Payload could see orchestration runtime credentials` warning. This is a positive check — absence of errors is not sufficient | | L9 | Work ran off-platform | Same run | `summarise_results` metadata shows `contributing_hosts` containing the payload pod names, not the run worker's hostname | | L10 | One workload dispatched per unit | Same run, or `uv run pytest -k dispatched_to_its_own` locally | `contributing_workers` has one entry per unit; the local test asserts four distinct external workers | | L11 | Payload **image** carries no orchestration dependency | `docker run --rm python -c "import importlib.util; print(importlib.util.find_spec('dagster') is not None)"` | Prints `False`. L1 proves the *source* does not import `dagster`; this proves the shipped image does not contain it either | Checks L4–L9 require a cluster. L1–L3, L10 and L11 run on a laptop and should gate every change to the payload or the dispatching op. `yaml/loosely-coupled/probe-pipes-k8s.yaml` clears L4–L9 in a single run from a throwaway namespace, without deploying a code location or a Dagster control plane. Prefer it over assembling the cluster checks by hand: the RBAC it grants is exactly the set L5 and L6 ask about, so a failure localises immediately. **L4–L10 were cleared on sandbox-cat-dat on 2026-08-31** using the sandbox variant `yaml/sandbox/probe-pipes-k8s-sandbox.yaml`. Run `cff9b348-bfc3-4ac1-ab51-a94892b8e3a0` reached `RUN_SUCCESS` in `dataprovider01`: four payload Jobs, four distinct payload pod hostnames in `contributing_hosts`, and the payload's `External payload started on …` lines in the dispatcher's log, which is the pod log stream doing its job as the message channel. No credentials were needed anywhere. One trap the run exposed. The platform's `dagster-svc-account` sets `automountServiceAccountToken: false`, and the Dagster chart overrides it to `true` on every pod it manages. A hand-written pod that does not is a plausible future failure: the pipes client selects in-cluster authentication correctly and then fails on `Service token file does not exist`, which points at Kubernetes rather than at the omission. --- ## 4. Common misconfiguration symptoms ### 4.1 Both setups | Symptom | Likely cause | Correction | |---|---|---| | Code location stuck in *Loading*, then errors | Entry point path in `codeServerArgs` does not match the image layout | Confirm `--python-file` matches `workspace.yaml`; both must be `src/distributed_execution/repository.py` | | `ImagePullBackOff` with `no match for platform` | Image published for a single non-matching architecture | Rebuild multi-arch with `docker buildx`, and pin a version tag rather than `latest` | | Run stays in `QUEUED` indefinitely | Run coordinator concurrency limit reached, or no schedulable node | Check `max_concurrent_runs` and tag concurrency limits; check node capacity and resource quota | | Run fails immediately with a serialisation error | Code location image and Dagster control-plane versions diverge | Align the `dagster` version in `pyproject.toml` with the chart's version and rebuild | ### 4.2 Tightly coupled | Symptom | Likely cause | Correction | |---|---|---| | Run pod starts, then fails with a connection timeout to port 5432 | Execution namespace NetworkPolicy does not permit egress to Postgres | Add an egress rule for the metadata database, or move run pods to an already-approved namespace via `jobNamespace` | | `env_vars_missing` is non-empty in `report_execution_target` metadata | Env vars are set on the code location deployment but not on the run pod | Add them under `runLauncher.config.k8sRunLauncher.runK8sConfig.containerConfig.env` — code location env is **not** inherited by run pods | | A literal `vault:...` string appears as a value at runtime | Vault mutating webhook did not process the pod | Verify the `vault.security.banzaicloud.io/*` annotations are on the **run pod** template, not only the code location pod | | Steps hang in `STARTING` with `k8s_job_executor` | Service account lacks Job create/watch permission | Apply `yaml/tightly-coupled/rbac-step-executor.yaml` and confirm with `kubectl auth can-i` | | `contributing_hosts` shows one host when `k8s_job_executor` is configured | Run tags or Launchpad config overrode the executor, or the image predates the change | Confirm the code location reloaded after the image bump; check the run's *Config* tab for an `execution:` override | | Step pods `OOMKilled` under fan-out | Per-step memory limit applied per pod, aggregate exceeded quota | Raise `step_k8s_config` limits or lower step concurrency; the two multiply | | Postgres refuses connections once fan-out grows | Each step pod is an independent DB client | Reduce step concurrency, raise the Postgres connection limit, or move the fan-out step to a loosely coupled target | ### 4.3 Loosely coupled | Symptom | Likely cause | Correction | |---|---|---| | Op fails with `No pipes messages received from the external payload` | The message path is broken, not the workload | Work through L6 then L4. The payload very likely ran and succeeded; only its reporting was lost | | Payload pod `Completed`, but Dagster shows no payload log lines | A log shipper is intercepting or truncating stdout | Exclude the payload namespace from the shipper, or switch to an object-storage message reader | | Op hangs until `pod_wait_timeout` (default 24 h) | Payload Job never scheduled — quota, node selector or image pull | Check `kubectl -n describe job `; lower `pod_wait_timeout` so the failure surfaces quickly | | `403 Forbidden` creating the Job | Dispatcher service account lacks Job create permission | Apply `yaml/loosely-coupled/rbac-pipes-dispatch.yaml` in the **payload** namespace | | Warning: `Payload could see orchestration runtime credentials` | Payload pod inherited run-pod env or a Vault annotation | Remove the inherited env; the payload should receive only what `extras` and explicit `env` pass it | | Payload exits non-zero but the run reports success | Exit status not being checked, or messages read before failure | Confirm the dispatching op returns through `_result_from_pipes`; do not swallow `PipesClientCompletedInvocation` errors | | Payload receives no `units` | `extras` key mismatch between dispatcher and `pipes.get_extra()` | Both sides must use the same key; a typo yields a `KeyError` inside the payload | | Run cancelled in the UI, payload pod keeps running | Cancellation is not propagated to dispatched workloads automatically | `delete_pod_on_completion` handles the normal path; for cancellation, verify orphaned Jobs and add a cleanup sensor | > **Partly observed.** The pipes rows were exercised on sandbox-cat-dat on > 2026-08-31. Rows describing the tightly coupled `k8s_job_executor` still follow > from the dagster-k8s API rather than from observation. > > One symptom the cluster run added, absent from the table above: a service > account with `automountServiceAccountToken: false` — which the platform's > `dagster-svc-account` uses — makes the pipes client fail with > `ConfigException: Service token file does not exist`. It reads as a Kubernetes > fault; the fix is `automountServiceAccountToken: true` on the pod. --- ## 5. Evidence retention For each workflow's first execution, attach to the workflow's repository or change record: 1. The run ID and its final status. 2. The `report_execution_target` output metadata block (pod identity, namespace, env var presence). 3. The `summarise_results` metadata block (`contributing_hosts` and `contributing_workers`), which proves which execution target was actually used and that the fan-out reached it. 4. For `k8s_job_executor`, the output of `kubectl get jobs -l dagster/run-id=`. 5. For the loosely coupled target, the payload image digest and the run log line emitted by `pipes.log` — together they prove which payload version ran and that the message channel was open. Items 2 and 3 together are sufficient to demonstrate that the configured execution target is the one that ran — which is the point of the checklist.