Eureka Moments

When the Best Model Is the One You Stop Trying to Make Smarter

A travel brand needed an updated marketing mix model and an interactive planning tool. The most useful version turned out to be the most conceptually humble one.

Most stories about model rebuilds celebrate the final spec. This one is about why the simplest spec won — and why the three more sophisticated versions we tried first didn't.

The brief was familiar: a national travel brand with a long-cycle, multi-channel media mix wanted to refresh its marketing mix model on the latest data, add two new channels that had recently come online, and replace a static PDF report with an interactive scenario planner the marketing team could actually use. Two and a half years of weekly booking and spend data across hundreds of local markets, ten paid channels, and a mix of mail, video, audio, search, and social.

The Wrong Question

The instinctive question when sitting down with a richer dataset is: how much more model can the data now support? Per-market channel effects? Per-segment lift curves? A Bayesian hierarchy with partial pooling? With several years of weekly data and hundreds of geographies, the parameter budget feels generous.

The honest answer turned out to be: not nearly as much as it looks. The data was rich in some dimensions and structurally thin in others. Long-cycle purchase decisions (weeks to months from impression to booking) meant that weekly-grain models tried to fit a precise time-lag that doesn't really exist in the data. Aggregating to monthly fixed the lag problem but reduced the panel to fewer than 30 national observations. And across those observations, the brand had a habit — sensible in business terms, deadly in econometric terms — of ramping all its channels together when it wanted to push bookings. Highly correlated regressors plus a short panel is a recipe for unstable channel attribution.

The signature of an over-specified model is that channel effects flip sign or swing by half whenever you add or drop a single channel. That isn't sampling noise. That's the data telling you it can't tell the channels apart.

Four Attempts

We worked through the ladder of complexity from the top down. Each attempt was a reasonable thing to try. Each one failed for a structural reason worth naming.

FOUR ATTEMPTS, ONE THAT WORKED 1. Free channel effects per local market, weekly Hundreds of markets × ten channels × weekly noise. Parameters wouldn't identify; attribution swung between markets without real-world meaning. REJECTED 2. Pooled across markets with per-market intercepts, monthly Better. But a recently launched channel with only a few months of data absorbed an outsized share of an unrelated booking rebound. REJECTED 3. National monthly, free channel effects Fewer than 30 monthly observations against ten correlated channels. The estimator zeroed out roughly half the channels and piled everything onto one. REJECTED 4. National monthly, prior-driven, one scalar fit Channel effectiveness held fixed from the prior version. The data fits one dial: are those priors collectively right? Stable, interpretable, honest. SHIPPED
The first three attempts each added something the data couldn't support. The fourth subtracted everything the data couldn't say.

The Spec That Worked

The version that shipped looks almost embarrassingly simple. Channel effectiveness — the cost per incremental booking for each channel — was held fixed at values established by the previous model. The new fit estimated exactly one media-related parameter: a single scalar that scales all channel contributions together. Below one means the priors were collectively overstated; above one means understated. Everything else in the model is baseline structure — intercept, seasonality, holiday effects, promotional flags.

The result was a model with stable, interpretable channel contributions that pass every "drop a channel and refit" check. The fit explained roughly four-fifths of monthly booking variation, with single-digit mean error, on a panel that the more sophisticated models couldn't pin down.

The honest move when the data can't identify what you want is to stop trying to identify it. Fix what you trust; fit only what the data can say.

The Other Half of the Project

Stakeholders don't spend their time staring at residual plots. They spend it in the planner: sliding spend up and down, swapping diminishing-returns assumptions, dragging flight schedules across the calendar, looking at projected bookings and return-on-ad-spend in real time.

We built the planner as a single self-contained HTML file with thirteen tabs across three sections — a scenario builder, a set of geographic and seasonal explorers, and a model-diagnostics view. No backend, no login server, no deploy pipeline. The whole tool — model outputs, panel data, interactive charts, the saturation math — lives inside a file that can be opened from a Dropbox folder or pushed to a basic web host. A password gate handles the casual-sharing concern without pretending to be real authentication.

The planner does its own forward math in the browser: applies Hill saturation per channel, computes projected bookings and return-on-ad-spend live as sliders move, and shows the result side-by-side with a frozen baseline so every change reads as a delta from "what we did last year." That live feedback loop — change a number, see the answer — is what makes the difference between a model that gets cited in slides and a tool that actually shapes a media plan.

Two Lessons

The first lesson is the modeling one: when adding model features makes the results less stable, the answer isn't a different set of features. It's fewer of them. Modern modeling literature is full of techniques for squeezing more signal out of thin data — partial pooling, regularization, Bayesian priors — and they're all correct in their place. But the most honest move under genuine data scarcity is to commit to what you trust and only ask the data what it can answer.

The second is about where value actually lands. The model deserves rigor, but the tool deserves equal attention — because the tool is what stakeholders interact with, week after week. An hour spent making a planner slider feel right moves more decisions than an hour spent on a fourth round of model selection. Future projects of this shape should budget accordingly.


Working with a measurement problem where the data is rich in some dimensions and thin in others? Say hello.