Predicted incrementality with PIE
PIEModel predicts the incrementality of a campaign that has not been tested,
using a BART regression fitted on campaigns whose incrementality an experiment
measured. It follows Gordon, Moakler and Zettelmeyer (2026), NBER Working Paper
35044, but uses Bayesian Additive Regression Trees, which give a posterior
for each prediction where the paper’s random forest gives a point estimate
(src/ammm/pie/model.py:80). The module is alpha, so its API and defaults may
change. Install it with the pie extra.
When PIE is appropriate
PIE suits an organisation that has run a corpus of comparable experiments (geo tests, conversion-lift or ghost-ad holdouts) and needs an estimate for campaigns that cannot each be tested. Its predictions transfer the experiments to new campaigns, so they are valid only if the corpus represents the campaigns being predicted and the recorded features explain the variation in incrementality. A few dozen experiments spanning the target campaigns is a practical minimum; with fewer, the posterior intervals are wide and the held-out check below has little power.
Data contract
Each row of X is one experiment and y holds its measured incrementality,
in units you choose and apply consistently, such as incremental conversions
per unit of spend. The model reads only the columns it is configured to use and
ignores the others (src/ammm/pie/model.py:302).
| Argument | Role |
|---|---|
pre_determined_features | Columns known before launch, such as objective, vertical, planned budget or audience type. Always used. |
post_determined_features | Columns observed only after launch, such as exposure rate, click-through rate or last-click conversions. Used only when use_post_determined_features=True. |
use_post_determined_features | True by default. Set it to False for predictions made before launch; the model then neither trains on nor requires the post-launch columns. |
standard_error_column | Optional column of per-experiment standard errors in the units of y. |
target_column | Name of the target in the PyMC graph and saved inference data. It does not select a column from X. |
Non-numeric columns are label-encoded and split one level at a time.
Categories absent from training raise an error at prediction time. Missing or
infinite feature values raise an error at fit and prediction time because
PIEModel does not impute; drop or fill those rows and record how
(src/ammm/pie/model.py:279). A numeric prediction input outside the training
range raises a UserWarning, because BART holds its prediction constant beyond
the edge of the corpus (src/ammm/pie/model.py:465).
Choose the feature boundary from the decision time. Post-launch features often carry most of the predictive signal, but a prediction made before launch cannot use them. The paper also splits the sample within each campaign so that post-launch features are not computed from the same outcomes as the measured lift; AMMM4 does not implement that split, so a post-launch feature derived from the experiment’s own outcome can overstate accuracy.
Measurement error
Supplying standard_error_column stops the model from treating every
experiment as exact. The likelihood for experiment i becomes
Normal(f(x_i), sqrt(sigma**2 + se_i**2)), which marginalises the latent true
incrementality, so a precise experiment carries more weight than a noisy one
(src/ammm/pie/model.py:437). Predictions set the standard error to zero, so
their draws describe true incrementality rather than a new noisy measurement
(src/ammm/pie/model.py:497). In a synthetic check where half the corpus
measured 1.0 with standard error 0.01 and half measured 0.0 with standard error
5.0, the mean prediction was 0.51 without the column and 0.96 with it
(tests/pie/test_model.py, test_standard_errors_downweight_noisy_experiments).
Fit a pre-launch model
The following synthetic example fits a pre-launch model with standard errors.
import numpy as np
import pandas as pd
from ammm.pie import PIEModel
rng = np.random.default_rng(1)
n = 120
corpus = pd.DataFrame({
"objective": rng.choice(["conversions", "traffic"], size=n),
"planned_budget": rng.uniform(5_000, 100_000, size=n),
"audience_size": rng.uniform(1e5, 5e6, size=n),
"lift_se": rng.uniform(0.02, 0.2, size=n),
})
true_lift = (
0.4 + 0.3 * (corpus["objective"] == "conversions")
+ corpus["planned_budget"] / 2e5
)
lift = pd.Series(rng.normal(true_lift, corpus["lift_se"]), name="lift")
def make_model():
return PIEModel(
pre_determined_features=["objective", "planned_budget", "audience_size"],
post_determined_features=[],
use_post_determined_features=False,
standard_error_column="lift_se",
target_column="lift",
model_config={"bart": {"m": 50, "alpha": 0.95, "beta": 2.0}},
)
pie = make_model()
pie.fit(corpus, lift, draws=500, tune=500, chains=2, random_seed=42)
The default configuration uses 200 trees; this example uses 50 to keep the run
short. Check the usual sampler diagnostics before reading the predictions.
PGBART does not mix in the way NUTS does, so an R-hat above 1.01 on the bart
variable is common; run more draws and chains and compare predictions between
chains rather than relying on R-hat alone.
Predict and read the uncertainty
Prediction does not need the standard-error column, and it replaces the model’s feature data, so run the diagnostics in the next section first.
new_campaigns = corpus.drop(columns="lift_se").head(5)
draws = pie.predict_posterior(new_campaigns, extend_idata=False)
summary = draws.quantile([0.05, 0.5, 0.95], dim="sample")
The 5 to 95 percent interval from predict_posterior reflects the uncertainty
in the BART function given the corpus. It excludes three sources of error:
whether the corpus represents the predicted campaign, whether the features
capture what drives incrementality, and any bias in the experiments
themselves. Treat a narrow interval for a campaign unlike the corpus as a
symptom of extrapolation rather than as evidence of precision.
Feature diagnostics with pymc-bart
pymc-bart already provides interpretation tools that accept a fitted
PIEModel. Run them immediately after fit and before any prediction. The
feature matrix is the encoded training data, and
the labels follow the model’s feature order.
import pymc_bart as pmb
X_train = pie.model["X"].get_value()
labels = list(pie.model.coords["feature"])
vi = pmb.compute_variable_importance(
pie.idata, pie.model["bart"], X_train,
model=pie.model, method="VI", random_seed=42,
)
pmb.plot_variable_importance(vi, labels=labels)
pmb.plot_pdp(
pie.model["bart"], X_train, xs_interval="quantiles",
var_idx=[labels.index("planned_budget")], random_seed=42,
)
pmb.plot_variable_inclusion(pie.idata, X_train, labels=labels)
Variable importance shows how much predictive accuracy is lost when each feature is removed, whereas variable inclusion counts how often trees split on a feature, which is cheaper but less reliable when features are correlated. Partial dependence plots show the marginal predicted response to one feature. All three describe the fitted predictor; none shows that a feature causes a change in incrementality, and partial dependence for a categorical feature is on its encoded integer scale.
Held-out coverage check
The most useful acceptance check is whether held-out experiments fall inside their predicted intervals at the stated rate. The following K-fold loop refits on four fifths of the corpus, predicts the remaining fifth, adds back each held-out experiment’s measurement error and records whether the measured lift falls inside the 90 percent interval.
from sklearn.model_selection import KFold
covered = []
for train_idx, test_idx in KFold(5, shuffle=True, random_state=0).split(corpus):
fold = make_model()
fold.fit(corpus.iloc[train_idx], lift.iloc[train_idx],
draws=300, tune=300, chains=2, random_seed=42)
held_out = corpus.iloc[test_idx]
tau = fold.predict_posterior(held_out, extend_idata=False).values
se = held_out["lift_se"].to_numpy()[:, None]
measured_draws = tau + rng.normal(0.0, se, size=tau.shape)
low, high = np.quantile(measured_draws, [0.05, 0.95], axis=1)
measured = lift.iloc[test_idx].to_numpy()
covered.append((measured >= low) & (measured <= high))
coverage = np.concatenate(covered).mean()
On this synthetic corpus the coverage was 0.93 against a nominal 0.90. Coverage
well below nominal on a real corpus means the intervals are too narrow, which
is typically caused by missing features, unrecorded standard errors or
heterogeneity between experiment types. Random folds test interpolation within
the corpus. To test transfer to a new advertiser, vertical or market, hold out
whole groups with GroupKFold, because the paper’s cold-start diagnostics
are not implemented.
Limits of PIE predictions
A PIE prediction is a model estimate, not experimental evidence, and it is not
yet a suitable lift-test calibration input for an MMM. Lift-test calibration
needs a spend change, a response change and its standard error (x,
delta_x, delta_y and sigma) from a study of the modelled channel
(src/ammm/mmm/lift_test.py:384). A PIE prediction is learned from experiments
that may already calibrate the same MMM, so adding it would count that evidence
twice, and its interval omits the transfer error described above. Keep
measured experiments as the calibration input and report PIE predictions
alongside the MMM.
AMMM4 does not implement three parts of the paper:
- within-campaign sample splitting for post-launch features (paper section 4.2);
- extrapolation and cold-start diagnostics across advertiser segments (paper section 5.3);
- the decision framework comparing go and no-go choices with experiment-based decisions (paper section 6).
For the other optional integrations, see PIE and optional integrations.
Implementation reference: src/ammm/pie/model.py:80, src/ammm/mmm/lift_test.py:384.