Fit and evaluate inference

Select inference settings for the quantities you need to estimate, then inspect both computation and model fit. A fixed draw count cannot guarantee reliable tail probabilities, channel rankings or budget comparisons.

Inference methods

model.fit accepts mcmc, map, demz, advi and fullrank_advi. Use MCMC for the documented multi-chain diagnostic workflow. MAP is a point estimate; variational inference approximates the posterior family. Diagnostics that require chains cannot be inferred from a point estimate or an optimisation completion flag.

The Python quickstart selects nuts_sampler="pymc". Alternative sampler backends require installed dependencies and may retain different diagnostic fields. Use explicit seeds and retain settings for comparisons.

Computation

Inspect divergences, rank-normalised R-hat, bulk and tail effective sample size (ESS), Monte Carlo standard error, trace behaviour and sampler-specific energy and tree-depth information. R-hat near one is necessary evidence of mixing, not proof of convergence; ESS requirements depend on the precision needed for the particular estimand. Consult the diagnostic definitions.

Investigate divergences and poor mixing through the parameterisation, priors, data and model geometry. More draws do not repair a wrong likelihood or weak identification. A higher target acceptance can help a numerical integration problem but does not establish the model is appropriate.

Continue the Python quickstart:

import arviz as az

summary = az.summary(model.idata, var_names=["saturation_beta", "y_sigma"])
print(summary)
model.plot.diagnostics.posterior_predictive()

Prediction and comparison

Posterior predictive checks compare replicated observations with features that matter: levels, variability, extremes, temporal structure and subgroup patterns. An in-sample match can coexist with poor forecasts or implausible attribution.

Use time-respecting splits. The runner holdout provides one fresh aggregate fit; TimeSliceCrossValidator provides a Python rolling-origin workflow. Continue with X, y and an ordinary model specification:

from ammm.mmm.time_slice_cross_validation import TimeSliceCrossValidator

cv = TimeSliceCrossValidator(n_init=32, forecast_horizon=8, date_column="date", step_size=8)
cv_results = cv.run(
    X, y, mmm=model, original_scale_vars=["y"],
    sampler_config={"draws": 500, "tune": 500, "chains": 2, "cores": 1,
                    "random_seed": 42, "nuts_sampler": "pymc"},
)
prediction_table = cv.summary.predictions()

The explicit original_scale_vars=["y"] registers the response required by cv.summary.predictions() on each new fold graph. Each fold needs its own diagnostic review. Do not carry target-fitted features, preprocessing statistics or calibration outcomes from its test period into training. The generic CV interface is not a guarantee of leakage-free custom preprocessing.

LOO and WAIC are predictive criteria conditional on their observation unit and likelihood factorisation. Compare compatible models on the same outcomes and scale; review uncertainty and influential observations. Large Pareto-k values indicate unstable importance sampling. Pointwise leave-one-out is not a substitute for future-period validation in dependent time series. Review diagnostic gate policy separately from the substantive predictive and causal assessment.

Respond to failed diagnostics

Use troubleshooting to connect divergences, poor mixing and unstable predictive comparisons to specific investigations. Inspect each fold or sensitivity fit separately. FE likelihood terms score within contrasts and CRE terms score collapsed units, so do not rank their raw LOO/WAIC values against ordinary observation-level scores without a compatible predictive target (src/ammm/mmm/fixed_effects.py:377, src/ammm/mmm/correlated_random_effects.py:542).

Implementation reference at 7cb7f20: src/ammm/mmm/mmm.py:1731, src/ammm/mmm/time_slice_cross_validation.py:57.