Graham McNicoll
text published 2026-06-25 · Open on LinkedIn ↗
Teams routinely misunderstand averages in experimentation and what those averages are hiding. I see product teams treat the average effect as the decision. Variant B beat A, p-value valid, so the product was shipped. An average is a summary, not an explanation. Inside that 3% you can have one segment with a strong positive response and three segments that were neutral or hurt by the change. The net looks green. The experience for most of your users got worse. You have to watch out for the Simpson's Paradox. The aggregate tells one story. The segmented view tells the opposite. Running the experiment is not the hard part. Interrogating the data is. The questions you care about are why it worked and for whom. See your results sliced by segment. Start at growthbook.io.