Methodologists are Statisticians. Right? Right?

Confession

I’m often introduced in meetings as “our resident statistician” and invariably I look around the room fearful that the jig was up and I would be called out as a sham and a charlatan. In my mind, a STATISTICIAN is someone who holds advanced degrees in Statistics or Mathematical Statistics, not someone who, like me, has done courses in advanced statistics or econometrics as required coursework in a different discipline such as Economics or academic Marketing. Of course, this distinction is well, academic in most business settings and my training and experience are often all that are necessary for the problems at hand.

That said, a recent experience led me to think about the differences between the way an honest-to-goodness Statistician would frame certain issues and the way applied statisticians such as myself frame the same problem.

To Test or Not to Test

A client recently asked a straightforward question. They had run a study on a new medical device surveying both HCPs and patients. Patients split naturally into two groups: those with prior experience using an injectable biologic, and those who had never used one. HCPs, meanwhile, answered on behalf of both patient types. The client wanted to know whether it was appropriate to statistically test these two groups against each other.

The rote answer is easy. Since each HCP is rating both patient types, that comparison calls for a test designed for paired responses – commonly called a dependent-samples paired test. Since the patient groups are made up of different individuals, that comparison calls for a standard test for independent groups. Both are well established, and the difference between them comes down entirely to how the data was collected, not to anything more interesting. I crafted my email accordingly and, as is customary these days, asked Claude to review it for clarity and accuracy. Claude gave me a Pass (phew!) and then suggested that I should add a cautionary note about the danger of stat-testing several attributes simultaneously – essentially, the possibility that many will turn up “significant” by chance alone, even if nothing real is going on.

I did not quite see what the concern was - in my mind, I had already assumed that Claude was conflating the issue with multiple paired-groups comparisons – I said as much.

The Statistician’s answer

Statistically, (so spake Claude) the concern is real and has a name: the multiple comparisons problem, or family-wise error rate inflation. Run one test at the standard 0.05 threshold and, if there's truly no effect, you have a 5% chance of a false positive. Run twenty independent tests under the same conditions, and the odds that at least one comes back significant by chance climb to roughly 64%. Academic fields have formal tools for this: Bonferroni corrections, false discovery rate control, pre-registered primary endpoints. This isn't controversial. If you're willing to interpret a scattering of significant results across a large attribute battery as a coherent finding, some of what you're looking at is just noise wearing the clothing of insight.

In over thirty years of doing market research, this is not something I have ever spent sleepless nights about. None of my mentors (illustrious as they are) taught me to worry about this problem. Why is that? Surely all those Econ and Marketing PhDs couldn’t all be ignorant of this issue.

Tukey’s HSD Anyone?

At this point, most market research methodologists are probably up in arms, yelling Tukey, or Bonferroni at the screen. And they would be right – we do account for multiple comparisons but along ONE dimension - banner point testing, the standard practice of comparing many groups or segments against each other across a crosstab, routinely applies exactly this kind of correction. Tukey's HSD, Duncan's multiple range test, Bonferroni-adjusted letter testing, these are standard tools built specifically to control the error rate that comes from comparing many groups pairwise. Anyone who has run a banner test with more than a handful of columns has used one of these, whether or not they thought of it in those terms.

So, (continued Claude), what's missing isn't the concept. It's the other axis. The correction gets applied when the multiplicity comes from comparing many groups to each other. It essentially never gets applied when the multiplicity comes from comparing two groups across many attributes, which is the exact same statistical mechanism, just running along a different dimension of the data. Ask around the industry about testing two groups across a twenty-item attribute battery, and multiple comparisons corrections rarely enter the conversation. And this is the crux - I have never once heard a client or colleague raise it, even though many of those same researchers correct for multiplicity every time they run a banner test.

Domain Knowledge as Shield

Back to the question then – are we practitioners fatally flawed in our understanding of how statistics works? Not quite. Here's the informal check that substitutes for a formal correction. Chance findings tend to scatter randomly across a battery of attributes. They don't cluster around a coherent story, because there's no underlying mechanism generating them. Real findings, by contrast, tend to cohere with what you already know about the groups being compared.

Take the biologic-experienced versus biologic-naive comparison. If the attributes that come back significant are things like injection site comfort, needle anxiety, and confidence in self-administration, that's exactly the cluster you'd predict from lived experience with an injectable. The pattern hangs together. A chance-driven false positive, by contrast, would be as likely to show up on an attribute like device color preference as on comfort with injection, because chance has no opinion about which attributes matter. When a pattern of significant results lines up with domain logic, that coherence is itself evidence against a chance explanation, not just an interesting footnote.

This is worth taking seriously as a real statistical argument, not merely a researcher's intuition dressed up after the fact. It's closer to a Bayesian argument than people usually give it credit for. A finding that fits your prior is more likely to be real; a finding with no plausible mechanism behind it deserves more scrutiny even at the same p-value. Experienced researchers are, in effect, running an informal prior-plausibility filter on top of the formal test, and that filter works well for the most part.

Post-hoc Rationalization (Shield at 5%)  

The domain knowledge argument is what I used to push back against Claude’s assertions. Fair enough (said Claude). However, the trouble is that this filter only works when it's applied honestly, and it's easy to apply it dishonestly without noticing.

It works well when the researcher has a clear, pre-existing view of what a plausible pattern should look like and checks the results against that view. It breaks down when the “plausible story” gets constructed after the fact, built to explain whichever handful of attributes happened to come back significant. That's confirmation bias operating under the cover of domain expertise, and it looks identical to legitimate pattern-checking from the outside. The only real difference is whether the expectation existed before you saw the results or was assembled afterward to fit them.

It also breaks down as the attribute battery grows. A plausibility check on three or four attributes tied to a specific hypothesis is a meaningful filter. A plausibility check on forty attributes, where you're implicitly reasoning “well, three or four of these seem like they'd make sense,” starts to lose its power, because with a battery that large, a coherent-sounding story is easy to construct around whatever happens to be significant, chance included.

Borrowing from the Clinicians

Clinical trials solved a version of this problem (continued Claude) by distinguishing primary endpoints from secondary ones. A small number of outcomes are specified in advance as the ones the study is actually designed to answer, and those get the full statistical rigor, corrections included where appropriate. Everything else collected along the way is exploratory: interesting, worth reporting, but explicitly not the basis for a confident claim.

Market research would benefit from borrowing that discipline more explicitly, even informally. Before running a group comparison across a large attribute battery, name the small number of attributes that are actually tied to the business question at hand, the ones a stakeholder would act on directly. Treat those as confirmatory and give the pattern-check the weight it deserves there. Treat everything else in the battery as exploratory, worth a look, worth flagging if something jumps out, but not something to build a messaging strategy around without further validation.

The value of this isn't that it slows anything down. It's that it makes explicit a distinction that experienced researchers are often already making instinctively and puts it in a form a client or a junior analyst can actually see and evaluate, rather than leaving it as an unstated judgment call buried in the analyst's head.

Verdict

The absence of formal multiple comparisons corrections in market research isn't sloppiness - it reflects a different control mechanism, plausibility based on domain knowledge, that does real statistical work when applied honestly and up front. The problem isn't that this substitute exists. It's that it is applied implicitly, which makes it impossible to tell whether it's being applied well or simply invoked after the fact to bless whatever the data happened to produce. Naming it turns an invisible judgment call into a deliberate practice, one that can be done well or poorly, rather than something that just happens.

So, are Methodologists Statisticians?

Yes, albeit a custom-built version, one designed to specifically address the kinds of problems we face in Market Research and (occasionally) in econometric data. To keep things honest, let’s just stick with Methodologist, shall we?

Next
Next

Incorporating a Naïve Bayes Classifier in a Choice Simulator