[SIMPL-30451] Distributed execution reference service and container-cluster configuration guide #1
@@ -43,12 +43,12 @@ location image.
|
||||
### 2.1 Sandbox state, 2026-08-31
|
||||
|
||||
The T checks were assessed against sandbox-cat-dat by reading the live
|
||||
`dagster-instance` ConfigMap in `dataprovider01`. None of them could be *executed*:
|
||||
the tightly coupled Kubernetes path needs this code location registered on the
|
||||
platform Dagster, which is blocked on the control plane skew, and unlike the
|
||||
pipes transport it cannot be exercised from a standalone pod — every step pod
|
||||
connects to the metadata database on its own, so they must share instance
|
||||
storage rather than a local SQLite file.
|
||||
`dagster-instance` ConfigMap in `dataprovider01` and by running T1 and T4 from a
|
||||
pod there. T6 and T7 could not be executed: the tightly coupled Kubernetes path
|
||||
needs this code location registered on the platform Dagster, which is blocked on
|
||||
the control plane skew, and unlike the pipes transport it cannot be exercised
|
||||
from a self-contained pod — every step pod connects to the metadata database on
|
||||
its own, so they must share instance storage rather than a local SQLite file.
|
||||
|
||||
What the configuration shows:
|
||||
|
||||
@@ -60,6 +60,46 @@ What the configuration shows:
|
||||
| T5 | Passes. `dagster-role`, bound to `dagster-svc-account`, grants `batch/jobs` with create, delete, get, list, patch, update and watch |
|
||||
| T8 | `fail_pod_on_run_failure: true` is set |
|
||||
|
||||
**T1 and T4 were executed on 2026-08-31** from a pod in `dataprovider01`
|
||||
(`yaml/sandbox/probe-tightly-coupled-sandbox.yaml`), using a local SQLite
|
||||
instance rather than the platform's Postgres.
|
||||
|
||||
| # | Result |
|
||||
|---|---|
|
||||
| T1 | Reachability **passes** — `check_tcp_reachable('pg-cluster.common01.svc.cluster.local', 5432)` returned `passed: True`. The env-var half of the check does not hold; see below |
|
||||
| T4 | **Passes.** `tightly_coupled_local_job` reached `RUN_SUCCESS` with `contributing_hosts: ['distexec-tc-probe-hbdfl']` — one entry — and four distinct PIDs in `contributing_workers`. Separate processes, same machine, exactly as the row predicts |
|
||||
|
||||
`tightly_coupled_in_process_job` also succeeded, reporting a single worker.
|
||||
|
||||
### T1's env-var check is wrong for this platform
|
||||
|
||||
The run emitted:
|
||||
|
||||
> Orchestration runtime env vars not visible to this process:
|
||||
> `DAGSTER_POSTGRES_HOST`, `DAGSTER_POSTGRES_USER`, `DAGSTER_POSTGRES_DB`.
|
||||
> Expected for a loosely coupled target; a misconfiguration for a tightly coupled one.
|
||||
|
||||
That warning would fire on a **correctly configured** platform run pod. The chart
|
||||
does not set those variables: it injects `DAGSTER_PG_PASSWORD` from
|
||||
`dagster-postgresql-secret` and bakes host, user and database into `dagster.yaml`.
|
||||
So `TIGHTLY_COUPLED_ENV_VARS` in `preflight.py` names variables Simpl does not
|
||||
use, and the check reports a misconfiguration that is not one. The reachability
|
||||
probe above is the part that carries real evidence.
|
||||
|
||||
### Why T6 and T7 remain unexecuted
|
||||
|
||||
The run made the obstacle concrete. Step outputs went to
|
||||
`/dagster-home/storage/<run-id>/…` via `PickledObjectFilesystemIOManager`. Under
|
||||
`k8s_job_executor` each step is its own pod with its own volume, so a downstream
|
||||
step cannot read an upstream output. Running it needs both a shared I/O manager —
|
||||
S3, or the RWX `dagster-shared-pvc` the launcher already mounts at
|
||||
`/dagster/shared` — and shared instance storage, which means Postgres.
|
||||
|
||||
The only Postgres available is the platform's live `dataprovider01_dagster`, and
|
||||
these images carry dagster 1.13.19 against a 1.12.8 control plane. Pointing a
|
||||
newer Dagster at an older schema is not a safe thing to try on a shared database.
|
||||
Pin dagster to the control plane version and rebuild first.
|
||||
|
||||
**C6 passes, contrary to what `job_namespace` alone suggests.** The launcher's
|
||||
`job_namespace: dagster` is only a default. Each code location ships a
|
||||
`DAGSTER_CONTAINER_CONTEXT` environment variable — generated by the chart — whose
|
||||
|
||||
99
yaml/sandbox/probe-tightly-coupled-sandbox.yaml
Normal file
99
yaml/sandbox/probe-tightly-coupled-sandbox.yaml
Normal file
@@ -0,0 +1,99 @@
|
||||
# Tightly coupled probe for sandbox-cat-dat - the part that is safe to run.
|
||||
#
|
||||
# Covers T4 and proves the run pod can reach the metadata database, using a
|
||||
# LOCAL SQLite instance. It deliberately does NOT write to the platform's
|
||||
# Postgres. See "What this cannot cover" below.
|
||||
#
|
||||
# Runs three things in one pod:
|
||||
# 1. A TCP probe to Postgres using this repository's own check_tcp_reachable.
|
||||
# 2. tightly_coupled_in_process_job - single process, no fan-out.
|
||||
# 3. tightly_coupled_local_job - multiprocess fan-out inside one pod (T4).
|
||||
#
|
||||
# --- What this cannot cover ------------------------------------------------
|
||||
#
|
||||
# T6 and T7 need `tightly_coupled_k8s_job`, and that cannot use a local SQLite
|
||||
# instance: k8s_job_executor puts every step in its own pod, and each step pod
|
||||
# opens the metadata database itself. That is the definition of tightly coupled.
|
||||
# The step pods must therefore share instance storage, which means Postgres.
|
||||
#
|
||||
# The only Postgres available is `dataprovider01_dagster`, the platform's live
|
||||
# metadata database, and these images carry dagster 1.13.19 against a control
|
||||
# plane on 1.12.8. Pointing a newer Dagster at an older schema is not a safe
|
||||
# operation on a shared database: at best it refuses to start with a schema
|
||||
# mismatch, and it is not worth finding out the worse case on someone else's
|
||||
# instance. Pin dagster to the control plane version and rebuild before trying.
|
||||
#
|
||||
# T1 is only partly covered. Its stated evidence expects DAGSTER_POSTGRES_HOST,
|
||||
# DAGSTER_POSTGRES_USER and DAGSTER_POSTGRES_DB in the pod environment; the
|
||||
# platform sets none of them - it passes DAGSTER_PG_PASSWORD and bakes the rest
|
||||
# into dagster.yaml. Left unset here deliberately, so the run reports the real
|
||||
# deployment's state rather than a flattering one.
|
||||
#
|
||||
# --- Running ---------------------------------------------------------------
|
||||
#
|
||||
# kubectl -n dataprovider01 apply -f yaml/sandbox/probe-tightly-coupled-sandbox.yaml
|
||||
# kubectl -n dataprovider01 logs -f job/distexec-tc-probe
|
||||
# kubectl -n dataprovider01 delete job distexec-tc-probe
|
||||
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: distexec-tc-probe
|
||||
namespace: dataprovider01
|
||||
labels:
|
||||
app.kubernetes.io/name: distributed-execution
|
||||
app.kubernetes.io/component: probe
|
||||
spec:
|
||||
backoffLimit: 0
|
||||
activeDeadlineSeconds: 1800
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app.kubernetes.io/name: distributed-execution
|
||||
app.kubernetes.io/component: probe
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
serviceAccountName: dagster-svc-account
|
||||
automountServiceAccountToken: true
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: runner
|
||||
image: gitea.dataprovider01.sandbox-cat-dat.simpl-europe.eu/j.r/distributed-execution:5122da4691f9
|
||||
imagePullPolicy: IfNotPresent
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop:
|
||||
- ALL
|
||||
env:
|
||||
- name: DAGSTER_HOME
|
||||
value: /dagster-home
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
echo "===== T1: can this pod reach the metadata database? ====="
|
||||
python -c "from distributed_execution.preflight import check_tcp_reachable; print(check_tcp_reachable('pg-cluster.common01.svc.cluster.local', 5432))"
|
||||
echo "===== tightly_coupled_in_process_job ====="
|
||||
dagster job execute -f src/distributed_execution/repository.py -j tightly_coupled_in_process_job
|
||||
echo "===== T4: tightly_coupled_local_job ====="
|
||||
dagster job execute -f src/distributed_execution/repository.py -j tightly_coupled_local_job
|
||||
volumeMounts:
|
||||
- name: dagster-home
|
||||
mountPath: /dagster-home
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 512Mi
|
||||
limits:
|
||||
cpu: "1"
|
||||
memory: 1Gi
|
||||
volumes:
|
||||
- name: dagster-home
|
||||
emptyDir: {}
|
||||
Reference in New Issue
Block a user