Run cff9b348-bfc3-4ac1-ab51-a94892b8e3a0 of loosely_coupled_k8s_job reached RUN_SUCCESS in dataprovider01 on sandbox-cat-dat: four payload Jobs, four distinct payload pod hostnames in contributing_hosts, and the payload log lines in the run log. Checks L4-L10 are now observed rather than derived. Torn down afterwards; the namespace was left as it was found. Corrects two claims that were wrong. The registry does NOT require a pull secret - a bare GET returns 401, but that is the opening move of the Docker token handshake, and completing it anonymously returns the manifest for both images. And PIPES_PAYLOAD_NAMESPACE was set to a namespace named dagster, which does not exist on that cluster; the release is called dagster but runs in dataprovider01. Adds the sandbox probe variant. The generic probe cannot run there: sandbox access is Rancher project-scoped, so a new namespace grants its creator nothing and Role creation is denied everywhere. It is not needed either, since dagster-role already carries the exact pipes permissions. Changelog: added
13 KiB
Distributed Execution: Readiness Checklist
Pre-run validation for a workflow's execution target. Companion to the user guide; this document covers AC3.
Run these checks before the first execution of a workflow, and again after any change to the run launcher, the executor, the target namespace or the code location image.
L4–L10 were cleared on sandbox-cat-dat on 2026-08-31. Checks L1–L3 and L11, which exercise the payload contract, the message-parsing path and image-level isolation, are verified locally. What remains unverified on a cluster is the tightly coupled Kubernetes path — the T rows and the Kubernetes rows of section 4.3 — which uses
k8s_job_executorrather than the pipes client.
1. Common checks — both setups
| # | Check | How to verify | Expected evidence |
|---|---|---|---|
| C1 | Code location loads | Dagster UI → Deployment → Code locations | Location distributed-execution shows status Loaded, with a recent load timestamp and no error banner |
| C2 | Jobs are registered | Dagster UI → Jobs | The jobs listed in the guide's section 5.5 appear under the code location |
| C3 | Image tag matches the intended release | kubectl -n dagster get deploy -l dagster/code-location=distributed-execution -o jsonpath='{.items[*].spec.template.spec.containers[*].image}' |
Tag equals the version in pipeline.variables.sh; never latest |
| C4 | Image architecture matches the nodes | docker manifest inspect <image> |
Includes linux/amd64; a manifest with only linux/arm64 produces no match for platform at pull time |
| C5 | Run launcher type is as intended | kubectl -n dagster get cm dagster-instance -o yaml |
run_launcher block shows K8sRunLauncher and the expected job_namespace |
| C6 | Target namespace exists and is schedulable | kubectl get ns <namespace> and kubectl -n <namespace> get resourcequota |
Namespace is Active; remaining quota exceeds the job's aggregate requests |
2. Tightly coupled checks
| # | Check | How to verify | Expected evidence |
|---|---|---|---|
| T1 | Run pod reaches the metadata database | Launch tightly_coupled_in_process_job |
Run reaches SUCCESS; report_execution_target output metadata lists DAGSTER_POSTGRES_HOST, DAGSTER_POSTGRES_USER and DAGSTER_POSTGRES_DB under env_vars_present, and env_vars_missing is empty |
| T2 | Vault injection works | Same run; inspect the run pod | kubectl -n dagster describe pod <run-pod> shows the vault-env init container completed; no vault: literal remains in the process environment |
| T3 | Object storage is reachable | Same run, if the workflow uses S3 | No EndpointConnectionError in run logs; S3_ENDPOINT_URL present in env_vars_present |
| T4 | Multiprocess fan-out actually fans out | Launch tightly_coupled_local_job |
Run succeeds; summarise_results metadata shows one entry per unit in contributing_workers and a single entry in contributing_hosts — separate processes, same machine |
| T5 | RBAC permits step Jobs | kubectl -n dagster auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev |
Returns yes; required only for k8s_job_executor |
| T6 | Step pods are actually created | Launch tightly_coupled_k8s_job, then kubectl -n dagster get jobs -l dagster/run-id=<run-id> |
One Job per mapped unit; contributing_hosts now shows one entry per unit, not one |
| T7 | Step pod egress is permitted | Same run | Steps do not hang in STARTING; run logs contain no connection timeouts to port 5432 |
| T8 | Failure surfaces as a pod failure | Force a step failure in a scratch namespace | failPodOnRunFailure: true is set, and the step pod reports Failed rather than Completed |
3. Loosely coupled checks
| # | Check | How to verify | Expected evidence |
|---|---|---|---|
| L1 | Payload contract is intact | uv run pytest tests/test_loosely_coupled.py |
test_payload_does_not_import_dagster passes — the payload imports dagster_pipes only |
| L2 | Pipes round trip works | Launch loosely_coupled_subprocess_job |
Run reaches SUCCESS; run logs contain the payload's External payload started on … line, proving pipes.log crossed the channel |
| L3 | Silence is treated as failure | Same test module | test_silent_message_path_is_treated_as_failure passes — an empty message list raises rather than yielding an empty result |
| L4 | Payload image is pullable by the target cluster | kubectl -n <payload-ns> run pull-probe --image=<payload-image> --restart=Never --command -- true |
Pod reaches Completed; no ImagePullBackOff. Both Gitea images pull anonymously — a bare registry GET returns 401, but that is the start of the Docker token handshake, not a refusal |
| L5 | Dispatcher can create Jobs in the payload namespace | kubectl -n <payload-ns> auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev |
Returns yes |
| L6 | Dispatcher can read pod logs — the message channel | kubectl -n <payload-ns> auth can-i get pods/log --as=system:serviceaccount:dagster:dagster-dev |
Returns yes. A no here breaks reporting without failing the workload |
| L7 | Payload Job is actually created | Launch loosely_coupled_k8s_job, then kubectl -n <payload-ns> get jobs -l app.kubernetes.io/name=distributed-execution-payload |
One Job per dispatch, labelled dagster/execution-target=loosely-coupled |
| L8 | Payload has no orchestration connectivity | Same run; inspect run logs | No Payload could see orchestration runtime credentials warning. This is a positive check — absence of errors is not sufficient |
| L9 | Work ran off-platform | Same run | summarise_results metadata shows contributing_hosts containing the payload pod names, not the run worker's hostname |
| L10 | One workload dispatched per unit | Same run, or uv run pytest -k dispatched_to_its_own locally |
contributing_workers has one entry per unit; the local test asserts four distinct external workers |
| L11 | Payload image carries no orchestration dependency | docker run --rm <payload-image> python -c "import importlib.util; print(importlib.util.find_spec('dagster') is not None)" |
Prints False. L1 proves the source does not import dagster; this proves the shipped image does not contain it either |
Checks L4–L9 require a cluster. L1–L3, L10 and L11 run on a laptop and should gate every change to the payload or the dispatching op.
yaml/loosely-coupled/probe-pipes-k8s.yaml clears L4–L9 in a single run from a
throwaway namespace, without deploying a code location or a Dagster control
plane. Prefer it over assembling the cluster checks by hand: the RBAC it grants
is exactly the set L5 and L6 ask about, so a failure localises immediately.
L4–L10 were cleared on sandbox-cat-dat on 2026-08-31 using the sandbox
variant yaml/sandbox/probe-pipes-k8s-sandbox.yaml. Run
cff9b348-bfc3-4ac1-ab51-a94892b8e3a0 reached RUN_SUCCESS in dataprovider01:
four payload Jobs, four distinct payload pod hostnames in contributing_hosts,
and the payload's External payload started on … lines in the dispatcher's log,
which is the pod log stream doing its job as the message channel. No credentials
were needed anywhere.
One trap the run exposed. The platform's dagster-svc-account sets
automountServiceAccountToken: false, and the Dagster chart overrides it to
true on every pod it manages. A hand-written pod that does not is a plausible
future failure: the pipes client selects in-cluster authentication correctly and
then fails on Service token file does not exist, which points at Kubernetes
rather than at the omission.
4. Common misconfiguration symptoms
4.1 Both setups
| Symptom | Likely cause | Correction |
|---|---|---|
| Code location stuck in Loading, then errors | Entry point path in codeServerArgs does not match the image layout |
Confirm --python-file matches workspace.yaml; both must be src/distributed_execution/repository.py |
ImagePullBackOff with no match for platform |
Image published for a single non-matching architecture | Rebuild multi-arch with docker buildx, and pin a version tag rather than latest |
Run stays in QUEUED indefinitely |
Run coordinator concurrency limit reached, or no schedulable node | Check max_concurrent_runs and tag concurrency limits; check node capacity and resource quota |
| Run fails immediately with a serialisation error | Code location image and Dagster control-plane versions diverge | Align the dagster version in pyproject.toml with the chart's version and rebuild |
4.2 Tightly coupled
| Symptom | Likely cause | Correction |
|---|---|---|
| Run pod starts, then fails with a connection timeout to port 5432 | Execution namespace NetworkPolicy does not permit egress to Postgres | Add an egress rule for the metadata database, or move run pods to an already-approved namespace via jobNamespace |
env_vars_missing is non-empty in report_execution_target metadata |
Env vars are set on the code location deployment but not on the run pod | Add them under runLauncher.config.k8sRunLauncher.runK8sConfig.containerConfig.env — code location env is not inherited by run pods |
A literal vault:... string appears as a value at runtime |
Vault mutating webhook did not process the pod | Verify the vault.security.banzaicloud.io/* annotations are on the run pod template, not only the code location pod |
Steps hang in STARTING with k8s_job_executor |
Service account lacks Job create/watch permission | Apply yaml/tightly-coupled/rbac-step-executor.yaml and confirm with kubectl auth can-i |
contributing_hosts shows one host when k8s_job_executor is configured |
Run tags or Launchpad config overrode the executor, or the image predates the change | Confirm the code location reloaded after the image bump; check the run's Config tab for an execution: override |
Step pods OOMKilled under fan-out |
Per-step memory limit applied per pod, aggregate exceeded quota | Raise step_k8s_config limits or lower step concurrency; the two multiply |
| Postgres refuses connections once fan-out grows | Each step pod is an independent DB client | Reduce step concurrency, raise the Postgres connection limit, or move the fan-out step to a loosely coupled target |
4.3 Loosely coupled
| Symptom | Likely cause | Correction |
|---|---|---|
Op fails with No pipes messages received from the external payload |
The message path is broken, not the workload | Work through L6 then L4. The payload very likely ran and succeeded; only its reporting was lost |
Payload pod Completed, but Dagster shows no payload log lines |
A log shipper is intercepting or truncating stdout | Exclude the payload namespace from the shipper, or switch to an object-storage message reader |
Op hangs until pod_wait_timeout (default 24 h) |
Payload Job never scheduled — quota, node selector or image pull | Check kubectl -n <payload-ns> describe job <name>; lower pod_wait_timeout so the failure surfaces quickly |
403 Forbidden creating the Job |
Dispatcher service account lacks Job create permission | Apply yaml/loosely-coupled/rbac-pipes-dispatch.yaml in the payload namespace |
Warning: Payload could see orchestration runtime credentials |
Payload pod inherited run-pod env or a Vault annotation | Remove the inherited env; the payload should receive only what extras and explicit env pass it |
| Payload exits non-zero but the run reports success | Exit status not being checked, or messages read before failure | Confirm the dispatching op returns through _result_from_pipes; do not swallow PipesClientCompletedInvocation errors |
Payload receives no units |
extras key mismatch between dispatcher and pipes.get_extra() |
Both sides must use the same key; a typo yields a KeyError inside the payload |
| Run cancelled in the UI, payload pod keeps running | Cancellation is not propagated to dispatched workloads automatically | delete_pod_on_completion handles the normal path; for cancellation, verify orphaned Jobs and add a cleanup sensor |
Partly observed. The pipes rows were exercised on sandbox-cat-dat on 2026-08-31. Rows describing the tightly coupled
k8s_job_executorstill follow from the dagster-k8s API rather than from observation.One symptom the cluster run added, absent from the table above: a service account with
automountServiceAccountToken: false— which the platform'sdagster-svc-accountuses — makes the pipes client fail withConfigException: Service token file does not exist. It reads as a Kubernetes fault; the fix isautomountServiceAccountToken: trueon the pod.
5. Evidence retention
For each workflow's first execution, attach to the workflow's repository or change record:
- The run ID and its final status.
- The
report_execution_targetoutput metadata block (pod identity, namespace, env var presence). - The
summarise_resultsmetadata block (contributing_hostsandcontributing_workers), which proves which execution target was actually used and that the fan-out reached it. - For
k8s_job_executor, the output ofkubectl get jobs -l dagster/run-id=<run-id>. - For the loosely coupled target, the payload image digest and the run log line
emitted by
pipes.log— together they prove which payload version ran and that the message channel was open.
Items 2 and 3 together are sufficient to demonstrate that the configured execution target is the one that ran — which is the point of the checklist.