Modelling contracts

Choose the model’s units, parameter sharing and response estimand explicitly. Ordinary MMM, FixedEffectsMMM and CorrelatedRandomEffectsMMM represent different statistical models; adding dims to an ordinary MMM does not select an econometric panel estimator.

Pooling

MMM(dims=("geo",)) indexes observations and parameters by geography. Independent prior entries with the same fixed hyperparameters are not a learned hierarchy. Sharing a parameter across geographies requires omitting that axis from its prior dimensions; partial pooling requires shared stochastic hyperparameters and an appropriate exchangeability assumption.

This self-contained example builds a partially pooled saturation amplitude:

import numpy as np
import pandas as pd
from pymc_extras.prior import Prior
from ammm.mmm import MMM, GeometricAdstock, LogisticSaturation

X = pd.DataFrame({
    "date": np.repeat(pd.date_range("2025-01-06", periods=4, freq="W-MON"), 2),
    "geo": ["north", "south"] * 4,
    "tv": [1., 10., 2., 20., 3., 30., 4., 40.],
})
y = np.array([10., 20., 12., 22., 14., 24., 16., 26.])
model = MMM(
    date_column="date", channel_columns=["tv"], dims=("geo",),
    adstock=GeometricAdstock(l_max=1), saturation=LogisticSaturation(),
    model_config={"saturation_beta": Prior(
        "LogNormal",
        mu=Prior("Normal", mu=0, sigma=1, dims="channel"),
        sigma=Prior("HalfNormal", sigma=0.5, dims="channel"),
        dims=("geo", "channel"),
    )},
)
model.build_model(X, y)

Only the configured amplitude uses this hierarchy. Other parameters retain their own dimensions and priors.

Scaling

The ordinary MMM default divides channels and target by their signed maxima over date and all panel dimensions. Channels retain separate scales; controls are not automatically standardised. For [-4, 2], the maximum is 2, not 4. Choose units before eliciting priors, and inspect the actual scalers.

Scaling dims are axes to reduce, with date always reduced. In a geo panel, dims=("geo",) pools the scaling reduction across geographies; dims=() keeps one target scale per geography and one channel scale per geography and channel. A dictionary configuration for separate scales is:

scaling = {
    "channel": {"method": "max", "dims": ()},
    "target": {"method": "max", "dims": ()},
}

DataDerivedScaling supports signed max and mean reductions. FixedScaling supplies explicit constants. Derived scales must be finite and non-zero, except that an all-zero media slice receives one. A zero target needs an explicit valid fixed scale and a likelihood admitting zero. Mean cancellation is rejected. LogSaturation forces channel scales to one, so it receives raw channel values. FE and CRE impose their own scaling contracts.

Response scale and likelihood

The default identity-link Normal model decomposes the conditional mean on the scaled target axis. Multiplying components by the target scale restores their units. An identity-link likelihood’s mu need not equal its expectation: truncation, for example, changes the mean. Therefore component sums need not reconcile to the posterior predictive mean for every supported likelihood.

The identity link rejects a conventional Prior("LogNormal", ...) because its mu is on the log scale. A LogNormalPrior instance is a separate supported special-prior interface, not a distribution name to pass into Prior. The log link supports the conventional LogNormal form. Check the likelihood’s support after scaling; choosing a familiar distribution name is insufficient.

LogNormal response and optimisation

The log link is experimental. For link="log", with LogNormal location mu and scale sigma, the conditional median is exp(mu) and the conditional mean is exp(mu + sigma**2 / 2). Multiply by the target scale for original units. A channel-removal contrast exp(mu) - exp(mu - component) is median-based, and individual removal contrasts generally do not sum to the all-channel contrast.

The default optimisation response is total_media_contribution_original_scale. For a log-link model it is a median-based contrast, even when averaged over posterior draws. compute_counterfactual_contributions_dataset(central_tendency="mean") changes that returned decomposition, not the optimiser’s objective.

A common positive mean-correction factor independent of allocation preserves a maximiser; factors varying across draws or cells can change weights and rankings. Choose and verify an explicit response variable when expected outcomes are the decision target. Custom effects may also require a different objective.

Calibration and causal interpretation

Calibration adds the likelihood of a supported experiment-derived quantity. Match its intervention, population, units, horizon and uncertainty, and justify transport to the modelling period. The lift API uses a static saturation contrast, rejects the log link, and does not replay arbitrary treatment and control exposure histories. Read the calibration guide.

A causal graph, partial pooling, successful inference and good predictions each provide different information. None alone establishes that a fitted response curve describes the proposed intervention. Use the identification guide and decision checks before interpreting allocations.

Verify a scale calculation

In the four-date panel in the pooling example, the default TV scale is 40 and the target scale is 26 across both geographies. With per-geography reductions, TV scales are 4 and 40 and target scales are 16 and 26. This follows the signed maximum reduction, not an inferred stochastic hierarchy (src/ammm/mmm/_mmm_init.py:167, src/ammm/mmm/scaling.py:537).

# Continue immediately after the pooling example above.
assert float(model.xarray_dataset._channel.max()) == 40.0
assert float(model.xarray_dataset._target.max()) == 26.0
print(model.model["channel_scale"].eval())
print(model.model["target_scale"].eval())

A prior amplitude of 0.5 on this pooled target scale corresponds to 13 outcome units before accounting for saturation and other multipliers. Changing the scale while retaining the same numerical prior therefore changes the substantive prior. For log-link quantities, if mu=0 and sigma=0.5, the scaled median is 1 and the mean is approximately 1.133; averaging median-based contrasts across posterior draws does not turn them into mean-based contrasts (src/ammm/mmm/link.py:622, src/ammm/mmm/link.py:693).

The retained YAML runner rejects log-link response preparation. Use the Python interface for that experimental link and verify the chosen response variable before optimisation (src/ammm/pipeline/stages/core.py:130).

Implementation reference at 7cb7f20: src/ammm/mmm/mmm.py:1375, src/ammm/mmm/scaling.py:123, src/ammm/mmm/link.py:210.