The Interaction One Channel Hides From the Other

Nested modeling in Marketing Mix Modeling measures each channel’s direct and indirect impact, so TV gets credit for the search it drives.

What this article argues
  • Why treating channels as independent quietly undercounts your upper-funnel media
  • How last-touch attribution hands TV’s credit to search
  • What endogeneity is, and why ordinary regression cannot fix it on its own
  • How two-stage least squares separates a channel’s direct and indirect impact
  • Where nested modeling fits alongside Bayesian MMM and incrementality tests

Most marketing mix models rest on a convenient assumption: that each channel works alone. TV drives sales, search drives sales, social drives sales, and none of them touch each other. It is a clean assumption, and it is wrong. Channels talk to each other constantly, and a model that ignores that conversation hands the credit to the wrong place.

This article sets out the technique MASS Analytics uses to fix it, called nested modeling, and is candid about the mechanics. It follows on from earlier MASS Analytics writing on the regression engine that still sits under modern MMM. Nested modeling is not a new engine; it is what you do when you respect one of that engine’s assumptions instead of pretending it holds. It sits in the same family as the other advanced techniques we have covered, including log-linear modeling and regional pooled regression.

Channels do not work in isolation

Endogeneity is not a TV problem. It shows up anywhere one marketing variable is partly driven by another. Display and social prime the branded search that later converts. Promotions lift baseline sales and pull in search traffic at once. Sponsorships and PR feed direct visits. In each case two variables the model treats as independent are quietly linked, and the moment that link exists, the assumption underneath ordinary regression breaks. The habit worth building is to look for it everywhere, not to file it under one channel pair.

Figure 1Direct and indirect impact is one of six behavioral effects a Marketing Mix Model has to represent. Nested modeling is the mechanic that separates it; the others each have their own.

The clearest illustration is still the classic one. A Telco runs a multichannel campaign. Some customers see the TV spot and buy directly, and any model gives that sale to TV. Others see the spot, search the brand a day later, click, and buy. Under last-touch, that sale is booked to search.

But the search never happened without the TV. The spot created the intent; search only harvested it. A model that treats TV and branded search as independent inputs books the assisted sale to search, so TV looks weaker than it is and the next budget round moves money the wrong way. Swap TV for any upper-funnel channel and the story repeats.

Endogeneity is not a TV problem. It appears wherever one channel quietly drives another, and the moment that link exists, ordinary regression starts crediting the wrong line.

Last-touch attribution pays the closer, not the setup

Think of a sales team. One rep spends weeks warming a prospect; another walks in on signing day. Pay commission only on the signature and you build a team of closers. No one is left doing the work that makes closing possible.

Naive MMM makes the same mistake at scale. It over credits lower-funnel channels that capture intent and under credits the upper-funnel channels that create it. The error is not random. It always favors the channels closest to the purchase, which are the easiest to cut before anyone notices the damage.

Ordinary least squares breaks when channels feed each other

The instinct is to reach for the standard tool. In regression analysis, coefficients are usually estimated with ordinary least squares, or OLS. It is unbiased on one condition: the explanatory variables must be independent of the model’s error term. Break that and the estimates are no longer trustworthy.

In the Telco case, the condition breaks. Search is not independent; it is partly driven by TV, and TV also drives sales directly. So search correlates with what the model has not captured, which lives in the error term. That correlation is endogeneity, and it makes OLS hand search a chunk of impact that belongs to TV. The coefficient on search is biased upward, and no amount of data cleaning fixes it.

Endogeneity is not a data-quality problem you can clean your way out of. It is structural, and ordinary least squares has no defense against it.

Nested modeling treats one channel as a mediator for another

Nested modeling names the relationship instead of assuming it away. You write two equations. The main equation is the one you care about: sales as a function of search, TV, and the rest. The secondary, or reduced, equation describes search as a function of TV and its own drivers.

That is a mediation structure. TV has a direct path to sales and an indirect path through search. Once both equations are written, you can ask what naive MMM cannot. How much of TV’s effect is direct, and how much travels through the search it provokes? A variable in the main model only has a direct impact. One in both has a direct impact plus an indirect one through the channel it feeds.

Figure 2. One common example: TV reaches sales directly and indirectly through the search it provokes. Nested modeling estimates both paths so upper-funnel media keeps its full credit. The same structure applies to any channel pair.

Two-stage least squares separates direct from indirect impact

The method that estimates this is two-stage least squares, or 2SLS. In the first stage you regress search on its exogenous drivers, TV among them, and keep the fitted values. These are the part of search explained by things outside the error term. In the second stage you drop those fitted values into the main sales equation in place of raw search.

That swap is the trick. Raw search was correlated with the error term, which biased the coefficient. Fitted search is built only from exogenous drivers, so the correlation is gone and the coefficient comes back unbiased. You recover both the direct and indirect impact of every channel. You can then adjust contribution and ROI to reflect what each one really did.

One honest caveat, because 2SLS is not magic. It only works if the first-stage drivers genuinely explain the mediating channel. If TV barely moves search in your data, you have a weak-instrument problem. The second stage then becomes unstable and can mislead you more than plain OLS. This is where the method stops being mechanical and needs a modeler who understands the business.

If your model says search is your best channel, ask what made people search. The answer is usually sitting in your upper-funnel budget, uncredited.

YouTube video
Nested modeling explained: how two-stage least squares separates a channel’s direct and indirect impact.

Where nested modeling sits next to Bayesian MMM and experiments

The move to Bayesian MMM has not made this go away. Bayesian frameworks handle uncertainty well and can carry interaction terms, but a mediation structure still has to be specified. If you do not tell the model that TV feeds search, it will not find the path on its own. You are back to independent channels with better error bars.

Experiments come at the effect from the other side. Run a geo holdout on TV and watch branded search move in the test markets. That is causal evidence the indirect path is real. That is the value of incrementality testing. Nested modeling is the observational counterpart. It quantifies the same split continuously, across the whole plan, and the strongest programs use both.

A Bayesian model will not discover that TV feeds search on its own. If you do not specify the mediation path, you are back to independent channels with better error bars.

Endogeneity is spreading, not shrinking

If anything, the problem is getting larger. Discovery is moving into AI answer engines. The branded search that nested modeling leans on as its mediating variable is migrating out of view. The consideration step now happens inside a ChatGPT or Perplexity conversation and resolves into direct traffic or a store visit. The signal that carried the upper funnel’s indirect effect is thinning as a result. The endogeneity did not go away; the variable that let us measure it is fading, which is harder. I have written about this as the paid search endogeneity problem in the age of AI search.

The takeaway reinforces the general point. Do not hard-code nested modeling as a TV-and-search trick. Treat it as the discipline of naming the real pathways in your mix. Keep adding new mediators as they appear, from AI citation share to direct traffic. That way the model keeps crediting demand to whatever created it.

Figure 3. A naive OLS model over-credits the channel closest to conversion; nesting returns the assisted credit to the channel that created the demand

What this changes in the next budget meeting

This is not only theory. In a published nested modeling case study, a retail brand ran heavy TV alongside branded search. Its first OLS model handed the uplift to the search line. A nested specification recovered the causal direction: branded search rose during TV windows, but that was TV’s demand surfacing downstream. The corrected credit led to a counterintuitive call: cutting branded search during TV peaks and lifting unbranded search. The result was a 12 per cent total sales uplift.

The broader payoff is a contribution table you can defend. Each upper-funnel channel then carries both its direct sales and the demand it pushed downstream. The top of the funnel stops looking like dead weight. Inside MassTer, the direct and indirect attribution view shows the direct, indirect, and total impact of every variable. The split is visible rather than buried in the math.

So before the next reallocation, ask one question of your own model. If it says search is your most efficient channel, does it know what made those people search? If not, you are not measuring efficiency. You are measuring who stood closest to the till.

Frequently asked questions about nested modeling

What is nested modeling in Marketing Mix Modeling?

Nested modeling is a technique that links two or more model equations so one marketing channel can act as a driver of another. It lets a Marketing Mix Model measure both the direct impact of a channel on sales and the indirect impact it has by driving other channels, such as TV creating the branded search that later converts.

Why can ordinary least squares not measure channel interactions?

Ordinary least squares assumes the explanatory variables are independent of the model’s error term. When one channel is partly driven by another, that channel correlates with unobserved factors in the error term, a condition called endogeneity. OLS then assigns biased coefficients, typically over crediting the downstream channel and under crediting the one that created the demand.

Does endogeneity only affect TV and paid search?

No. Endogeneity arises whenever one marketing variable is partly driven by another, so it appears across the mix: display and social driving branded search, promotions lifting both baseline and search, sponsorships feeding direct traffic, and increasingly AI answer engines shaping demand before any click. TV and branded search is the clearest example, not the only one, so the safe habit is to check every channel pair for a hidden dependency.

What is the difference between direct and indirect impact?

Direct impact is the effect a channel has on sales on its own. Indirect impact is the effect it has by influencing another channel that then drives sales. A TV campaign that sells directly and also lifts branded search has both a direct impact and an indirect impact routed through search. Total contribution is the sum of the two.

What is two-stage least squares?

Two-stage least squares, or 2SLS, is the estimation method behind nested models. In the first stage it predicts the mediating channel from its exogenous drivers. In the second stage it uses those fitted values in the main sales equation. Because the fitted values are not correlated with the error term, the resulting coefficients are unbiased, unlike a plain OLS estimate.

When should I use nested modeling instead of adding an interaction term?

Use nested modeling when one channel causally drives another and you want to split the second channel’s credit back to the first. An interaction term captures that two channels amplify each other but does not separate cause from effect or recover an indirect contribution. If the question is who created the demand, nested modeling answers it and a single interaction term does not.

Does nested modeling replace incrementality testing?

No. Incrementality testing gives causal evidence from a controlled experiment, such as a geo holdout on TV that moves branded search. Nested modeling estimates the same direct and indirect split observationally, across the whole media plan and continuously. The two are complementary: experiments validate the paths that the nested model quantifies at scale.

Key Takeaways
  • Treating channels as independent is the most common and most expensive assumption in MMM, because it systematically undercredits the upper-funnel media that creates demand.
  • Last-touch logic and naive regression make the same error: they pay the channel closest to the purchase and starve the one that set it up.
  • Endogeneity is why ordinary least squares cannot fix this on its own. When one channel drives another, its coefficient is biased and no amount of data cleaning removes the bias.
  • Nested modeling with two-stage least squares recovers both the direct and indirect impact of each channel, so contribution and ROI reflect the full effect of upper-funnel media.
  • The problem is intensifying as AI answer engines move demand off search, so treat nested modeling as a general discipline and keep adding new mediators as they appear.