Dr. Ramla Jarrar on the data required in marketing mix modeling, why collection and quality decide the result, and the check that protects the model before a single regression runs.
The Model Is Only As Honest As the Data You Hand It
Thin, wrongly specified, or never-questioned data compromises most marketing mix models that disappoint, before anyone even opens the modeling tool. The model then faithfully reproduces those flaws.
A marketing mix model reads everything that plausibly drove your sales over a historical period. That is its strength, and it is also the constraint people forget. The model can only ever be as honest as the data you hand it.
So the work that decides the result is not the choice of algorithm. It is the data you request, the metrics you pick to represent each channel, and the checks you run before a single variable reaches the model.
Data Collection Starts With a Request, Not a Download
The first real decision in an MMM project is not which model to run. It is figuring out the data required in marketing mix modeling, and asking for exactly that.
Collection begins with a formal data request, built from the business questions agreed at kickoff. Each item names three things: the category, whether TV, search, display, price, or distribution, the level of granularity, and the exact definition. For search, that means stating whether you want impressions, clicks, or both, because those are three different questions.
The discipline here is a tradeoff, not a maximization. More data buys more granular answers. Ask for too much and you burden the client team into missing deadlines and sending the wrong extract. The skill is requesting exactly the data the business questions require, and no more.
This is also where you set coherence. If the data does not map to the decisions the CMO will make, the CMO will read the outputs, file them, and never act on them.
The Metric Is a Claim, Not a Data Field
Spend is always available. It is rarely the right variable.
Every channel in a mix serves a purpose, and the same channel can serve a different purpose for two different companies. So the first question is never which metric is easiest. It is what you put that channel in the plan to do.
Take a year where Meta budget rises 20 percent. If impressions do not rise 20 percent, you did not reach 20 percent more people. Model that channel on spend and you will expect a sales lift the exposures never justified. Model it on impressions and the variable tells the truth. In a media market where unit costs have inflated, that gap is not a rounding error. It is the difference in the coefficient, and the coefficient is the difference in the investment recommendation.
My approach is to pull every metric a channel offers, impressions, clicks, views, spend, and choose the one that best captures what you meant that channel to achieve. For branded search I lean to clicks, for generic search to impressions. Spend is the common denominator, and sometimes it is genuinely the best choice. Easier just does not mean better.
MMM Can Only Answer What Your Data Supports
MMM cannot read what your data does not contain. This sounds obvious and is routinely ignored in scoping.
If you have never changed your prices, because that is how your category works, no model will hand you a price elasticity. A channel with almost no variance or almost no activity will be very hard to measure. Without a history of investment in a channel, econometrics cannot manufacture a read on it, however much you want one.
So set expectations against the data you actually hold. If you want to know which creative or which campaign worked, you need data at the campaign level. A daypart buying change needs data by daypart to show its impact. Tracking a competitor’s prices is the only way to know how much their pricing steals from you.
The productive move is to stop asking the model for what the data cannot support, and start building the data infrastructure that will support it next cycle. Investing in your data is the real lever on what MMM can do for you.
Data Preparation Is Where the Project Lives or Dies
Data collection and preparation is the most underestimated phase in the whole workflow. It consistently takes around 60 percent of total project time (MASS Analytics Benchmark Database). The challenge is almost never a shortage of data. It is the heterogeneity of it: different formats, different granularities, different time windows, all of which you have to harmonize before a single variable exists.
The best analysts I have known treat this phase as a key success factor, not a chore. Early in my career I worked with a lead who printed every variation of every dependent variable against every independent variable, pre and post transformation, and spent days in what she called her yellow folder, working out what the data was actually telling her. She was almost always right about which variables would make sense to the model.
Automation and AI have compressed weeks of this work into days, and they should. But the instinct to interrogate the data has not become optional. When people rush past preparation to reach the equation, they end up believing whatever the model says, with no due diligence on whether the variables and the coefficients make business sense.
Keep the spirit of enquiry. Use the new tools to run it faster.
Quality Is a Gate, Not a Review
A data check is a go or no-go decision made the moment data arrives, not a report you write later. Nothing moves to exploration until it passes.
Run every incoming feed through a fixed sequence: readability, completeness, consistency, duplication, granularity, anomaly detection, and structural breaks. All of them pass, or the feed goes back. On a modern platform this runs again on every refresh, so quality stays a discipline rather than a one-time effort.
One of MASS Analytics’ clients, one of the biggest furniture retailers in the USA, cut data preparation time by around 60 percent after automating this gate, reaching three model cycles a year where one used to be the norm.
Reconcile the KPI to the client’s own records before anything else. If your data file shows a sales total the finance system does not recognize, you resolve that in writing before sign-off. A model calibrated against a KPI the business does not recognize produces results the business will not trust.
When a project involves EU consumer data, source governance is part of the gate, not a footnote. MMM is privacy compliant by design because it works on aggregated data, but you still have to collect and transfer the feeds lawfully. Penalties under the GDPR reach 20 million euros or 4 percent of worldwide annual turnover, whichever is greater.
Never Assume the Data Is Correct
If I could leave an analyst with one instruction, it would be this. Never assume the data is correct.
Abundance is guaranteed at the start of a project. Quality is not. So cross-check before you ingest. When impressions, clicks, and spend arrive together, chart spend against clicks and confirm the ratios make sense before the numbers ever reach a model. When TV spend jumps 40 percent in a month but impressions fall, you have a question to resolve, a rate card change or an extraction error, not a variable to model as is.
Two written outputs make this defensible. A questions and status log, recording every query raised, its resolution state, and the client’s answer. And an anomalies and assumptions log, recording every unresolved anomaly and the assumption the team will carry into transformation. Together they are the audit trail. Every assumption is visible to anyone who reads the model six months later, when the debrief question finally comes.
An anomaly caught at the data check costs one client conversation. The same anomaly, caught after you build and present the model, can cost you the modeling phase, restarted.
Frequently Asked Questions
A marketing mix model needs four families of data over a common historical period: the KPI or dependent variable such as sales, subscriptions, or acquisitions, marketing activities such as price, promotions, distribution, and direct marketing, paid, owned, and earned media, and external factors such as seasonality and market trends. Collect the KPI first, at the most granular level available, even when the model runs at national level.
Use the metric that best represents what you meant the channel to do, not the one that is easiest to obtain. Spend is a fair default and works well in many cases, but it conflates media volume with unit cost. When a budget rises without a matching rise in impressions, a spend variable overstates activity. Pull every available metric per channel and choose per channel against the business objective.
As granular as the questions you need answered, and no more. To compare creatives or campaigns you need campaign-level data. To measure a daypart change you need data by daypart. Granularity you never collect is insight you can never recover, but granularity you do not need only burdens the data request and delays the project.
Data collection and preparation typically consume around 60 percent of total project time. The bottleneck is rarely volume. It is harmonizing sources that arrive in different formats, granularities, and time periods into one clean, model-ready dataset, and validating every feed before you use it. Automation compresses this work, but the analyst’s judgment on what the data is saying remains essential.
It is a go or no-go gate applied the moment data arrives, testing readability, completeness, consistency, duplication, granularity, anomalies, and structural breaks. It also includes reconciling the KPI to the client’s own finance records. Nothing passes to exploration until the check clears, and the same check runs on every subsequent data refresh.
Yes. MMM works on aggregated data, with no cookies, device identifiers, or individual-level tracking, so it is compliant with the GDPR and comparable frameworks by design. You still have to collect and transfer the data feeds lawfully, so confirming the legal basis for each EU data source belongs in the sign-off documentation.
What To Do Next
Before you question your next model’s result, question its inputs, starting with the data required in marketing mix modeling. Pull the data request for your current project and check three things. Does every variable carry the metric the business decision actually needs. Is the KPI reconciled to finance in writing. Is there a log somewhere that would let a stranger reconstruct every assumption in the model.
If any of those three is missing, the fix is not a better algorithm. It is a better data check, run before you touch the next feed.
The question to bring to your next planning meeting is simple. If the model told you to move a million in budget tomorrow, would you trust the data underneath it enough to sign.


