Diagnose a failed model workflow
Locate the failing operation and its evidence before changing the model or
sampler. These checks describe source commit
7cb7f20b357f23ee1683670953c7dd0092fa2084, reviewed on 1 October 2026.
A process exit, a diagnostic decision and a causal claim are separate results.
Locate the failure
For a retained runner invocation, open run_manifest.json, then the failing
stage’s error and 00_run_metadata/run.log; a traceback can also be retained as
00_run_metadata/error.traceback.txt. A failure before directory creation may
only have terminal output. --verbose includes library output and tracebacks
(src/ammm/pipeline/runner.py:156, src/ammm/pipeline/runner.py:295,
runme.py:288). See output schemas for stage states.
| Symptom | Check and remedy | Implementing boundary |
|---|---|---|
| Named target mismatch | Match the Series name to target_column, or intentionally use an unnamed Series; verify the outcome itself | src/ammm/mmm/_mmm_graph.py:42 |
| Non-finite, duplicate or incomplete labelled data | Inspect original date/panel keys and measurements; repair the data explicitly before building instead of filling outcomes with zero | src/ammm/mmm/data_conversion.py:189 |
| Zero/non-finite derived scale | Inspect signed-max/mean reduction and zero target; select defensible fixed scales or correct degenerate data | src/ammm/mmm/scaling.py:495 |
| Likelihood support error | Compare observed values after scaling with the configured likelihood support; log response requires positive targets | src/ammm/mmm/link.py:622 |
| Direct YAML builder cannot find data | Its data paths are process-relative; pass preloaded X/y or absolute paths. Runner data paths are YAML-relative | src/ammm/mmm/builders/yaml.py:137; src/ammm/pipeline/runner.py:356 |
y_original_scale absent from returned prior dataset | Inspect registered deterministics in model.idata["prior"]; the returned dataset extracts prior_predictive | src/ammm/model_builder.py:673 |
| Cost calibration says original-scale contribution is missing | Build, register channel_contribution, add calibration, then fit the same data | src/ammm/mmm/_mmm_calibration.py:230 |
| Prior-only graph cannot be fitted or saved | Rebuild with the real target, then reapply graph-bound calibration and fit | src/ammm/mmm/mmm.py:1709; src/ammm/mmm/persistence.py:62 |
| Unsupported saved format | Use its originating release or reconstruct the full specification and refit; do not disable checks to simulate migration | src/ammm/mmm/persistence.py:62 |
| FE prior plot shows within-unit contrasts | Expected for the contrast likelihood; compare observed contrasts on the same scale. Posterior predictions reconstruct levels. The former missing-date failure is repaired in the current worktree | src/ammm/mmm/plotting/diagnostics.py:206 |
| Blocked holdout validation excludes calibration | Remove calibration for evaluation, or disable validation for a calibrated fit. Dry runs require changing the YAML because they ignore stage-disable overrides | src/ammm/pipeline/stages/validation.py:28 |
| Log model rejected by runner | Current response preparation requires identity link; use the experimental Python contract deliberately | src/ammm/pipeline/stages/core.py:130 |
| Dry run succeeds but quick execution fails | Dry run ignores sampler/quick/stage overrides and later stage work; inspect actual resolved config and failing stage | runme.py:163 |
| Missing diagnostic values | Check inference method, sampler and retained groups; unavailable evidence is skipped rather than a pass | src/ammm/mmm/diagnostic_gates.py:507 |
| Provider call absent | Inspect local rules and invocation status; do_not_use, local-only settings or a cache hit can explain it | src/ammm/pipeline/stages/ai_advisor.py:181 |
Investigate sampling evidence
Use the resolved diagnostic policy and parameter-level summaries together.
The bundled demo warns at R-hat >= 1.01 and fails at >= 1.05; any divergence
fails its divergence gate. Bulk/tail ESS warn at <= 400 and fail at <= 200
(data/demo/diagnostic_gates.yml:9, src/ammm/mmm/diagnostic_gates.py:507).
These are screening thresholds; Monte Carlo precision must suit the decision.
| Evidence | Investigation | What the change cannot establish |
|---|---|---|
| Divergences | Inspect affected parameter regions, likelihood support, scales, hierarchical parameterisation and prior predictive behaviour; investigate numerical integration settings after checking the model | A higher target acceptance does not repair confounding or prove unbiased sampling |
| High R-hat or low ESS | Examine chains, weakly constrained parameters and redundant components; compare parameterisations and defensible priors before merely extending sampling | A longer chain cannot create identifying variation |
| Large Pareto-k or LOO warning | Inspect influential observations and compatible scored units; consider direct held-out evaluation for the actual prediction task | An unstable importance-sampling score cannot reliably rank models |
| Poor holdout error or coverage | Check split design, leakage, factual adstock history, support and temporal regime changes | Tuning repeatedly on the same tail does not preserve an independent final test |
| Different channel estimates with similar fit | Compare baseline, priors, transformed exposure and adjustment assumptions; report inconclusive channel estimates where appropriate | Selecting the preferred decomposition does not establish its causal interpretation |
See inference and comparison, identification and the Stan diagnostic reference. Retain the original failure and the reason for each change so a reviewer can separate a repaired computation from a changed statistical question.
Recheck the same contract
Rerun the operation that failed and inspect its outputs, then check downstream
contracts affected by the change. For a data repair, rebuild, reapply calibration
and refit; for a saved-state problem, verify labelled prediction after loading.
For a runner integration defect, a graph-only dry run is insufficient: the
repaired stage and dependent retained workflow need execution checks
(src/ammm/mmm/_mmm_graph.py:182, src/ammm/pipeline/runner.py:215).
Installation and missing-extra failures belong in the
installation guide.