Files
distributed-execution/documents/user-guide/readiness-checklist.md
ILay aa49420ca6 [SIMPL-30787] Document container-cluster configuration inputs
Add the cluster-configuration input reference the guide was missing: which inputs are mandatory, conditional or optional in each setup, a reusable minimum input set, and worked examples expressed in the format each setup already uses. No configuration schema or key naming is prescribed.

Retitle the subject sections so they no longer carry one story's AC numbering, and add a traceability section mapping SIMPL-30787 and SIMPL-30451 acceptance criteria onto them.

Changelog: added
2026-09-02 18:48:01 +02:00

23 KiB
Raw Blame History

Distributed Execution: Readiness Checklist

Pre-run validation for a workflow's execution target. Companion to the user guide; this document is the validation checklist required by SIMPL-30451 AC3 and SIMPL-30787 AC3. Section 1 holds the checks common to both setups, section 2 the tightly coupled checks and section 3 the loosely coupled ones; every item states the evidence that proves it.

Run these checks before the first execution of a workflow, and again after any change to the run launcher, the executor, the target namespace or the code location image.

Both Kubernetes paths are cluster-verified. L4L10 were cleared on sandbox-cat-dat on 2026-08-31. The tightly coupled k8s_job_executor path reached RUN_SUCCESS through the platform launcher on 2026-09-01. Checks L1L3 and L11, which exercise the payload contract, message parsing and image isolation, are also verified locally. See section 2.1.


1. Common checks — both setups

# Check How to verify Expected evidence
C1 Code location loads Dagster UI → DeploymentCode locations Location distributed-execution shows status Loaded, with a recent load timestamp and no error banner
C2 Jobs are registered Dagster UI → Jobs The jobs listed in the guide's section 5.5 appear under the code location
C3 Image tag matches the intended release kubectl -n dagster get deploy -l dagster/code-location=distributed-execution -o jsonpath='{.items[*].spec.template.spec.containers[*].image}' Tag equals the version in pipeline.variables.sh; never latest
C4 Image architecture matches the nodes docker manifest inspect <image> Includes linux/amd64; a manifest with only linux/arm64 produces no match for platform at pull time
C5 Run launcher type is as intended kubectl -n <namespace> get cm dagster-instance -o yaml run_launcher block shows K8sRunLauncher. Note that job_namespace here is only a default — a code location's DAGSTER_CONTAINER_CONTEXT overrides it, so read that too before concluding where run pods land
C6 Target namespace exists and is schedulable kubectl get ns <namespace> and kubectl -n <namespace> get resourcequota Namespace is Active; remaining quota exceeds the job's aggregate requests. Also confirm the launcher's instance_config_map, postgres_password_secret and any PVC volumes exist in the namespace run pods actually use — these references do not cross namespaces (see section 2.1)

2. Tightly coupled checks

# Check How to verify Expected evidence
T1 Run pod reaches the metadata database Launch tightly_coupled_in_process_job Run reaches SUCCESS; report_execution_target metadata lists DAGSTER_PG_PASSWORD under env_vars_present without exposing its value, and env_vars_missing is empty. Successful step startup is the connectivity proof because the step loads run state from Postgres before executing the op
T2 Vault injection works Same run; inspect the run pod kubectl -n dagster describe pod <run-pod> shows the vault-env init container completed; no vault: literal remains in the process environment
T3 Object storage is reachable Same run, if the workflow uses S3 No EndpointConnectionError in run logs; S3_ENDPOINT_URL present in env_vars_present
T4 Multiprocess fan-out actually fans out Launch tightly_coupled_local_job Run succeeds; summarise_results metadata shows one entry per unit in contributing_workers and a single entry in contributing_hosts — separate processes, same machine
T5 RBAC permits step Jobs kubectl -n dagster auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev Returns yes; required only for k8s_job_executor
T6 Step pods are actually created Launch tightly_coupled_k8s_job, then kubectl -n dagster get jobs -l dagster/run-id=<run-id> One Job per mapped unit; contributing_hosts now shows one entry per unit, not one
T7 Step pod egress is permitted Same run Steps do not hang in STARTING; run logs contain no connection timeouts to port 5432
T8 Failure surfaces as a pod failure Force a step failure in a scratch namespace failPodOnRunFailure: true is set, and the step pod reports Failed rather than Completed

2.1 Sandbox state, 2026-08-31

The T checks were assessed against sandbox-cat-dat by reading the live dagster-instance ConfigMap in dataprovider01 and by running T1, T4, T6 and T7 from pods there.

What the configuration shows:

# Finding
C5 run_launcher is K8sRunLauncher; job_namespace: dagster is the default, overridden per code location to dataprovider01
T2 Vault injection is configured — pod_template_spec_metadata carries the banzaicloud annotations with role sandbox-cat-dat-role
T3 S3_ENDPOINT_URL is https://s3.sandbox-cat-dat.simpl-europe.eu, with access keys injected from Vault
T5 Passes. dagster-role, bound to dagster-svc-account, grants batch/jobs with create, delete, get, list, patch, update and watch
T8 fail_pod_on_run_failure: true is set

T1 and T4 were executed on 2026-08-31 from a pod in dataprovider01 (yaml/sandbox/probe-tightly-coupled-sandbox.yaml), using a local SQLite instance rather than the platform's Postgres.

# Result
T1 Reachability passescheck_tcp_reachable('pg-cluster.common01.svc.cluster.local', 5432) returned passed: True. The env-var half of the check does not hold; see below
T4 Passes. tightly_coupled_local_job reached RUN_SUCCESS with contributing_hosts: ['distexec-tc-probe-hbdfl'] — one entry — and four distinct PIDs in contributing_workers. Separate processes, same machine, exactly as the row predicts

tightly_coupled_in_process_job also succeeded, reporting a single worker.

T1's env-var check and platform contract

The run emitted:

Orchestration runtime env vars not visible to this process: DAGSTER_POSTGRES_HOST, DAGSTER_POSTGRES_USER, DAGSTER_POSTGRES_DB. Expected for a loosely coupled target; a misconfiguration for a tightly coupled one.

That warning was a false positive on a correctly configured platform run pod. The chart does not set those variables: it injects DAGSTER_PG_PASSWORD from dagster-postgresql-secret and stores host, user and database in dagster.yaml. The reference check now tests for DAGSTER_PG_PASSWORD and reports only its name, never its value. Database reachability is evidenced by successful step startup: a tightly coupled step must load its run state from the configured Dagster instance before report_execution_target can execute.

T6 and T7, executed 2026-08-31

Running them needed three things the earlier attempt lacked: dagster pinned to 1.12.8 to match the control plane, so writing to the live dataprovider01_dagster could not trigger a migration; step outputs moved onto the RWX dagster-shared-pvc, because the default I/O manager writes them under $DAGSTER_HOME/storage and each step pod otherwise gets its own empty volume; and automount_service_account_token: true in step_k8s_config, because dagster-svc-account disables it by default. The probe is yaml/sandbox/probe-k8s-executor-sandbox.yaml; the run is d05c736a-8ec7-4e59-8cb0-9cb238db40dc.

# Result
T6 Passes. Six Jobs labelled dagster/job=tightly_coupled_k8s_job were created, one per step, each in its own pod
T7 Passes. No step hung in STARTING and no step reported a connection timeout to port 5432. Every step pod opened the platform Postgres itself — the outputs and inputs below were all recorded through it

Four steps executed and succeeded: generate_work_units (four dynamic outputs), report_execution_target, and process_work_unit[unit_0] and [unit_1], each loading its input from the shared volume that a different pod wrote.

The run then failed, and summarise_results never started. The cause is a property of the probe, not of the code or the platform:

Detected run worker status UNKNOWN: DagsterK8sUnrecoverableAPIError … jobs.batch "dagster-run-d05c736a-…" is forbidden: User "system:serviceaccount:dataprovider01: dagster-svc-account" cannot get resource "jobs/status" … in the namespace "dagster"

Run monitoring polls for a run worker Job named dagster-run-<run-id>. The probe starts the run with dagster job execute from a hand-rolled Job, so no such object exists; and with no launcher-supplied container context the monitor looked in the launcher's default namespace, dagster, where dagster-svc-account holds no rights. The 403 turned absent into unknown, and the monitor failed the run at its first poll. process_work_unit[unit_2] and [unit_3] had their pods started already; both logged Skipping step execution … since the run is in status DagsterRunStatus.FAILURE and exited 0 — which is why all six Jobs read Complete while only four steps ran.

A run submitted the normal way, through the webserver to K8sRunLauncher, gets a dagster-run-<run-id> Job in dataprovider01 and is not exposed to this. Closing T6/T7 end to end, including summarise_results, needs the code location registered on the platform Dagster.

T6 and T7 closed end to end, 2026-09-01

With the code location registered on the platform Dagster, tightly_coupled_k8s_job was launched from the webserver and reached RUN_SUCCESS as run 1d8cb167-9fb8-4a34-a036-789eac381e13. K8sRunLauncher created dagster-run-1d8cb167-… in dataprovider01, K8sStepHandler created one dagster-step-… Job per step, and every step — including summarise_results — completed. No hand-rolled probe, no run-monitoring 403.

The last thing in the way was where step outputs are written. Two earlier attempts failed on FileNotFoundError: /tmp/dagster/storage/<run-id>/generate_work_units/result/unit_0, with every step pod using its own container filesystem. Mounting dagster-shared-pvc into the step pods was necessary but not sufficient: setting DAGSTER_HOME through step_k8s_config does not take effect, because the Dagster chart already injects DAGSTER_HOME=/tmp/dagster and the appended second entry does not displace the first. The fix is to bypass DAGSTER_HOME and pin the I/O manager instead — fs_io_manager.configured({"base_dir": SHARED_IO_BASE_DIR}) on the job, with SHARED_IO_BASE_DIR under the mounted PVC. The successful run logs confirm it:

Handled output "result" using IO manager "io_manager" — /dagster/shared/distributed-execution/storage/1d8cb167-…/generate_work_units/result/unit_0

The platform run exposed the false-positive DAGSTER_POSTGRES_* warning described under T1 above. The reference now checks the chart's actual DAGSTER_PG_PASSWORD injection instead.

Note: report_execution_target reported namespace: <not-in-kubernetes> on a step pod that plainly was in Kubernetes. The check reads the service account namespace file, which is absent when the token is not automounted. It is a property of the probe, not of the run.

C6 passes, contrary to what job_namespace alone suggests. The launcher's job_namespace: dagster is only a default. Each code location ships a DAGSTER_CONTAINER_CONTEXT environment variable — generated by the chart — whose k8s.namespace is dataprovider01, and the code location's context is merged over the launcher's, so it wins (dagster_k8s/container_context.py: namespace = other.namespace if other.namespace else self.namespace). Run pods therefore land in dataprovider01, where dagster-instance, dagster-postgresql-secret and dagster-shared-pvc all exist.

Both deployed code locations carry it, so both would launch correctly:

Code location DAGSTER_CONTAINER_CONTEXT k8s.namespace
template-code-location dataprovider01
semaphore-ui dataprovider01

The successful run confirms that the code-location container context places run and step pods in dataprovider01 with the required namespace-local dependencies.

The residual risk is narrower than a broken namespace, but real: a code location that reaches the launcher without a container context namespace inherits dagster and would fail on three missing objects. The chart always sets it, so this bites only hand-rolled pods — the reason the C6 row asks you to check the launcher's namespace-local dependencies rather than just the namespace.

An earlier revision of this section claimed C6 failed and that this explained the absence of runs. That was wrong: it read job_namespace without accounting for the container-context override.

3. Loosely coupled checks

# Check How to verify Expected evidence
L1 Payload contract is intact uv run pytest tests/test_loosely_coupled.py test_payload_does_not_import_dagster passes — the payload imports dagster_pipes only
L2 Pipes round trip works Launch loosely_coupled_subprocess_job Run reaches SUCCESS; run logs contain the payload's External payload started on … line, proving pipes.log crossed the channel
L3 Silence is treated as failure Same test module test_silent_message_path_is_treated_as_failure passes — an empty message list raises rather than yielding an empty result
L4 Payload image is pullable by the target cluster kubectl -n <payload-ns> run pull-probe --image=<payload-image> --restart=Never --command -- true Pod reaches Completed; no ImagePullBackOff. Both Gitea images pull anonymously — a bare registry GET returns 401, but that is the start of the Docker token handshake, not a refusal
L5 Dispatcher can create Jobs in the payload namespace kubectl -n <payload-ns> auth can-i create jobs --as=system:serviceaccount:dagster:dagster-dev Returns yes
L6 Dispatcher can read pod logs — the message channel kubectl -n <payload-ns> auth can-i get pods/log --as=system:serviceaccount:dagster:dagster-dev Returns yes. A no here breaks reporting without failing the workload
L7 Payload Job is actually created Launch loosely_coupled_k8s_job, then kubectl -n <payload-ns> get jobs -l app.kubernetes.io/name=distributed-execution-payload One Job per dispatch, labelled dagster/execution-target=loosely-coupled
L8 Payload has no orchestration connectivity Same run; inspect run logs No Payload could see orchestration runtime credentials warning. This is a positive check — absence of errors is not sufficient
L9 Work ran off-platform Same run summarise_results metadata shows contributing_hosts containing the payload pod names, not the run worker's hostname
L10 One workload dispatched per unit Same run, or uv run pytest -k dispatched_to_its_own locally contributing_workers has one entry per unit; the local test asserts four distinct external workers
L11 Payload image carries no orchestration dependency docker run --rm <payload-image> python -c "import importlib.util; print(importlib.util.find_spec('dagster') is not None)" Prints False. L1 proves the source does not import dagster; this proves the shipped image does not contain it either

Checks L4L9 require a cluster. L1L3, L10 and L11 run on a laptop and should gate every change to the payload or the dispatching op.

yaml/loosely-coupled/probe-pipes-k8s.yaml clears L4L9 in a single run from a throwaway namespace, without deploying a code location or a Dagster control plane. Prefer it over assembling the cluster checks by hand: the RBAC it grants is exactly the set L5 and L6 ask about, so a failure localises immediately.

L4L10 were cleared on sandbox-cat-dat on 2026-08-31 using the sandbox variant yaml/sandbox/probe-pipes-k8s-sandbox.yaml. Run cff9b348-bfc3-4ac1-ab51-a94892b8e3a0 reached RUN_SUCCESS in dataprovider01: four payload Jobs, four distinct payload pod hostnames in contributing_hosts, and the payload's External payload started on … lines in the dispatcher's log, which is the pod log stream doing its job as the message channel. No credentials were needed anywhere.

One trap the run exposed. The platform's dagster-svc-account sets automountServiceAccountToken: false, and the Dagster chart overrides it to true on every pod it manages. A hand-written pod that does not is a plausible future failure: the pipes client selects in-cluster authentication correctly and then fails on Service token file does not exist, which points at Kubernetes rather than at the omission.


4. Common misconfiguration symptoms

4.1 Both setups

Symptom Likely cause Correction
Code location stuck in Loading, then errors Entry point path in codeServerArgs does not match the image layout Confirm --python-file matches workspace.yaml; both must be src/distributed_execution/repository.py
ImagePullBackOff with no match for platform Image published for a single non-matching architecture Rebuild multi-arch with docker buildx, and pin a version tag rather than latest
Run stays in QUEUED indefinitely Run coordinator concurrency limit reached, or no schedulable node Check max_concurrent_runs and tag concurrency limits; check node capacity and resource quota
Run fails immediately with a serialisation error Code location image and Dagster control-plane versions diverge Align the dagster version in pyproject.toml with the chart's version and rebuild

4.2 Tightly coupled

Symptom Likely cause Correction
Run pod starts, then fails with a connection timeout to port 5432 Execution namespace NetworkPolicy does not permit egress to Postgres Add an egress rule for the metadata database, or move run pods to an already-approved namespace via jobNamespace
env_vars_missing is non-empty in report_execution_target metadata Env vars are set on the code location deployment but not on the run pod Add them under runLauncher.config.k8sRunLauncher.runK8sConfig.containerConfig.env — code location env is not inherited by run pods
A literal vault:... string appears as a value at runtime Vault mutating webhook did not process the pod Verify the vault.security.banzaicloud.io/* annotations are on the run pod template, not only the code location pod
Steps hang in STARTING with k8s_job_executor Service account lacks Job create/watch permission Apply yaml/tightly-coupled/rbac-step-executor.yaml and confirm with kubectl auth can-i
FileNotFoundError loading a step input from /tmp/dagster/storage/... Step pods each use their own container filesystem for I/O manager output Mount an RWX volume into step pods via step_k8s_config and pin fs_io_manager's base_dir under it. Setting DAGSTER_HOME in step_k8s_config does not work — the chart's DAGSTER_HOME=/tmp/dagster is injected first and wins
contributing_hosts shows one host when k8s_job_executor is configured Run tags or Launchpad config overrode the executor, or the image predates the change Confirm the code location reloaded after the image bump; check the run's Config tab for an execution: override
Step pods OOMKilled under fan-out Per-step memory limit applied per pod, aggregate exceeded quota Raise step_k8s_config limits or lower step concurrency; the two multiply
Postgres refuses connections once fan-out grows Each step pod is an independent DB client Reduce step concurrency, raise the Postgres connection limit, or move the fan-out step to a loosely coupled target

4.3 Loosely coupled

Symptom Likely cause Correction
Op fails with No pipes messages received from the external payload The message path is broken, not the workload Work through L6 then L4. The payload very likely ran and succeeded; only its reporting was lost
Payload pod Completed, but Dagster shows no payload log lines A log shipper is intercepting or truncating stdout Exclude the payload namespace from the shipper, or switch to an object-storage message reader
Op hangs until pod_wait_timeout (default 24 h) Payload Job never scheduled — quota, node selector or image pull Check kubectl -n <payload-ns> describe job <name>; lower pod_wait_timeout so the failure surfaces quickly
403 Forbidden creating the Job Dispatcher service account lacks Job create permission Apply yaml/loosely-coupled/rbac-pipes-dispatch.yaml in the payload namespace
Warning: Payload could see orchestration runtime credentials Payload pod inherited run-pod env or a Vault annotation Remove the inherited env; the payload should receive only what extras and explicit env pass it
Payload exits non-zero but the run reports success Exit status not being checked, or messages read before failure Confirm the dispatching op returns through _result_from_pipes; do not swallow PipesClientCompletedInvocation errors
Payload receives no units extras key mismatch between dispatcher and pipes.get_extra() Both sides must use the same key; a typo yields a KeyError inside the payload
Run cancelled in the UI, payload pod keeps running Cancellation is not propagated to dispatched workloads automatically delete_pod_on_completion handles the normal path; for cancellation, verify orphaned Jobs and add a cleanup sensor

Partly observed. The pipes rows were exercised on sandbox-cat-dat on 2026-08-31. The tightly coupled k8s_job_executor rows were closed on 2026-09-01 by run 1d8cb167-9fb8-4a34-a036-789eac381e13; the remaining rows in the table above still follow from the dagster-k8s API rather than from observation.

One symptom the cluster run added, absent from the table above: a service account with automountServiceAccountToken: false — which the platform's dagster-svc-account uses — makes the pipes client fail with ConfigException: Service token file does not exist. It reads as a Kubernetes fault; the fix is automountServiceAccountToken: true on the pod.


5. Evidence retention

For each workflow's first execution, attach to the workflow's repository or change record:

  1. The run ID and its final status.
  2. The report_execution_target output metadata block (pod identity, namespace, env var presence).
  3. The summarise_results metadata block (contributing_hosts and contributing_workers), which proves which execution target was actually used and that the fan-out reached it.
  4. For k8s_job_executor, the output of kubectl get jobs -l dagster/run-id=<run-id>.
  5. For the loosely coupled target, the payload image digest and the run log line emitted by pipes.log — together they prove which payload version ran and that the message channel was open.

Items 2 and 3 together are sufficient to demonstrate that the configured execution target is the one that ran — which is the point of the checklist.