The Data a Marketing Mix Model Needs, and How to Vet It Before You Build

Dr. Ramla Jarrar on the data required in marketing mix modeling, why collection and quality decide the result, and the check that protects the model before a single regression runs.

What This Article Argues
  • Why the data request, not the data download, is where an MMM project is won or lost
  • The difference between a metric and a data field, and why the choice changes the ROI you report
  • The quality gate that decides go or no-go before any model is built
  • The two documentation habits that keep a model defensible months later

The Model Is Only As Honest As the Data You Hand It

Thin, wrongly specified, or never-questioned data compromises most marketing mix models that disappoint, before anyone even opens the modeling tool. The model then faithfully reproduces those flaws.

A marketing mix model reads everything that plausibly drove your sales over a historical period. That is its strength, and it is also the constraint people forget. The model can only ever be as honest as the data you hand it.

So the work that decides the result is not the choice of algorithm. It is the data you request, the metrics you pick to represent each channel, and the checks you run before a single variable reaches the model.

Data Collection Starts With a Request, Not a Download

The first real decision in an MMM project is not which model to run. It is figuring out the data required in marketing mix modeling, and asking for exactly that.

Collection begins with a formal data request, built from the business questions agreed at kickoff. Each item names three things: the category, whether TV, search, display, price, or distribution, the level of granularity, and the exact definition. For search, that means stating whether you want impressions, clicks, or both, because those are three different questions.

The discipline here is a tradeoff, not a maximization. More data buys more granular answers. Ask for too much and you burden the client team into missing deadlines and sending the wrong extract. The skill is requesting exactly the data the business questions require, and no more.

This is also where you set coherence. If the data does not map to the decisions the CMO will make, the CMO will read the outputs, file them, and never act on them.

A data request is a design decision. What you fail to ask for at kickoff, you cannot recover at the modeling stage.

YouTube video
Watch: Introduction to Data Collection, from the MASS Analytics MMM course.

The Metric Is a Claim, Not a Data Field

Spend is always available. It is rarely the right variable.

Every channel in a mix serves a purpose, and the same channel can serve a different purpose for two different companies. So the first question is never which metric is easiest. It is what you put that channel in the plan to do.

Take a year where Meta budget rises 20 percent. If impressions do not rise 20 percent, you did not reach 20 percent more people. Model that channel on spend and you will expect a sales lift the exposures never justified. Model it on impressions and the variable tells the truth. In a media market where unit costs have inflated, that gap is not a rounding error. It is the difference in the coefficient, and the coefficient is the difference in the investment recommendation.

My approach is to pull every metric a channel offers, impressions, clicks, views, spend, and choose the one that best captures what you meant that channel to achieve. For branded search I lean to clicks, for generic search to impressions. Spend is the common denominator, and sometimes it is genuinely the best choice. Easier just does not mean better.

Impressions measure exposure. Clicks measure engagement. Spend measures investment. A model built on the wrong metric gives a precise answer to a question the business never asked.

MMM Can Only Answer What Your Data Supports

MMM cannot read what your data does not contain. This sounds obvious and is routinely ignored in scoping.

If you have never changed your prices, because that is how your category works, no model will hand you a price elasticity. A channel with almost no variance or almost no activity will be very hard to measure. Without a history of investment in a channel, econometrics cannot manufacture a read on it, however much you want one.

So set expectations against the data you actually hold. If you want to know which creative or which campaign worked, you need data at the campaign level. A daypart buying change needs data by daypart to show its impact. Tracking a competitor’s prices is the only way to know how much their pricing steals from you.

The productive move is to stop asking the model for what the data cannot support, and start building the data infrastructure that will support it next cycle. Investing in your data is the real lever on what MMM can do for you.

Data Preparation Is Where the Project Lives or Dies

Data collection and preparation is the most underestimated phase in the whole workflow. It consistently takes around 60 percent of total project time (MASS Analytics Benchmark Database). The challenge is almost never a shortage of data. It is the heterogeneity of it: different formats, different granularities, different time windows, all of which you have to harmonize before a single variable exists.

The best analysts I have known treat this phase as a key success factor, not a chore. Early in my career I worked with a lead who printed every variation of every dependent variable against every independent variable, pre and post transformation, and spent days in what she called her yellow folder, working out what the data was actually telling her. She was almost always right about which variables would make sense to the model.

Automation and AI have compressed weeks of this work into days, and they should. But the instinct to interrogate the data has not become optional. When people rush past preparation to reach the equation, they end up believing whatever the model says, with no due diligence on whether the variables and the coefficients make business sense.

Keep the spirit of enquiry. Use the new tools to run it faster.

Quality Is a Gate, Not a Review

A data check is a go or no-go decision made the moment data arrives, not a report you write later. Nothing moves to exploration until it passes.

Run every incoming feed through a fixed sequence: readability, completeness, consistency, duplication, granularity, anomaly detection, and structural breaks. All of them pass, or the feed goes back. On a modern platform this runs again on every refresh, so quality stays a discipline rather than a one-time effort.

One of MASS Analytics’ clients, one of the biggest furniture retailers in the USA, cut data preparation time by around 60 percent after automating this gate, reaching three model cycles a year where one used to be the norm.

Reconcile the KPI to the client’s own records before anything else. If your data file shows a sales total the finance system does not recognize, you resolve that in writing before sign-off. A model calibrated against a KPI the business does not recognize produces results the business will not trust.

When a project involves EU consumer data, source governance is part of the gate, not a footnote. MMM is privacy compliant by design because it works on aggregated data, but you still have to collect and transfer the feeds lawfully. Penalties under the GDPR reach 20 million euros or 4 percent of worldwide annual turnover, whichever is greater.

Seven checks, one gate: readability, completeness, consistency, duplication, granularity, anomaly detection, and structural breaks. All seven pass before exploration begins, or the feed goes back.

Never Assume the Data Is Correct

If I could leave an analyst with one instruction, it would be this. Never assume the data is correct.

Abundance is guaranteed at the start of a project. Quality is not. So cross-check before you ingest. When impressions, clicks, and spend arrive together, chart spend against clicks and confirm the ratios make sense before the numbers ever reach a model. When TV spend jumps 40 percent in a month but impressions fall, you have a question to resolve, a rate card change or an extraction error, not a variable to model as is.

Two written outputs make this defensible. A questions and status log, recording every query raised, its resolution state, and the client’s answer. And an anomalies and assumptions log, recording every unresolved anomaly and the assumption the team will carry into transformation. Together they are the audit trail. Every assumption is visible to anyone who reads the model six months later, when the debrief question finally comes.

An anomaly caught at the data check costs one client conversation. The same anomaly, caught after you build and present the model, can cost you the modeling phase, restarted.

Anything you do not have data on, you will not be able to measure. No model recovers what collection never captured.

Frequently Asked Questions

What is the data required in marketing mix modeling?

A marketing mix model needs four families of data over a common historical period: the KPI or dependent variable such as sales, subscriptions, or acquisitions, marketing activities such as price, promotions, distribution, and direct marketing, paid, owned, and earned media, and external factors such as seasonality and market trends. Collect the KPI first, at the most granular level available, even when the model runs at national level.

Should I model on spend or impressions?

Use the metric that best represents what you meant the channel to do, not the one that is easiest to obtain. Spend is a fair default and works well in many cases, but it conflates media volume with unit cost. When a budget rises without a matching rise in impressions, a spend variable overstates activity. Pull every available metric per channel and choose per channel against the business objective.

How granular should MMM data be?

As granular as the questions you need answered, and no more. To compare creatives or campaigns you need campaign-level data. To measure a daypart change you need data by daypart. Granularity you never collect is insight you can never recover, but granularity you do not need only burdens the data request and delays the project.

Why does data preparation take so long in MMM?

Data collection and preparation typically consume around 60 percent of total project time. The bottleneck is rarely volume. It is harmonizing sources that arrive in different formats, granularities, and time periods into one clean, model-ready dataset, and validating every feed before you use it. Automation compresses this work, but the analyst’s judgment on what the data is saying remains essential.

What is a data quality check in MMM?

It is a go or no-go gate applied the moment data arrives, testing readability, completeness, consistency, duplication, granularity, anomalies, and structural breaks. It also includes reconciling the KPI to the client’s own finance records. Nothing passes to exploration until the check clears, and the same check runs on every subsequent data refresh.

Is marketing mix modeling GDPR compliant?

Yes. MMM works on aggregated data, with no cookies, device identifiers, or individual-level tracking, so it is compliant with the GDPR and comparable frameworks by design. You still have to collect and transfer the data feeds lawfully, so confirming the legal basis for each EU data source belongs in the sign-off documentation.

What To Do Next

Before you question your next model’s result, question its inputs, starting with the data required in marketing mix modeling. Pull the data request for your current project and check three things. Does every variable carry the metric the business decision actually needs. Is the KPI reconciled to finance in writing. Is there a log somewhere that would let a stranger reconstruct every assumption in the model.

If any of those three is missing, the fix is not a better algorithm. It is a better data check, run before you touch the next feed.

The question to bring to your next planning meeting is simple. If the model told you to move a million in budget tomorrow, would you trust the data underneath it enough to sign.

Key Takeaways
  • The data required in marketing mix modeling can only be as honest as what you collect, so collection and quality decide the result before modeling begins.
  • The metric is a claim about what a channel does, not a neutral data field, and choosing spend by default can distort the ROI you report.
  • You cannot measure what you never collected, so scope expectations against the data you hold and invest in the infrastructure for the answers you want next.
  • Treat data quality as a go or no-go gate with a fixed set of checks, not a review you write up afterward.
  • Keep a questions log and an anomalies log, because the assumption you fail to record is the one that undermines the debrief months later.