FeaturedHow Sport Is Governed: Bodies, Arbitration, and Integrity
Science

Reading Statistics in the News: P-Values, Risk, and Sample Size

Significance, confidence intervals, relative risk and screening accuracy are routinely misreported. A practical guide to the handful of statistical ideas that decide whether a health headline means anything.

Editorial Team
Team reviewing data on a whiteboard
Photo: You X Ventures · Unsplash License

What a p-value is actually measuring

A p-value answers one narrow question: if there were genuinely no effect, how often would data at least as extreme as the data observed arise purely through random variation? A small p-value indicates that the observed result would be unusual in a world where nothing was happening. That is all it says. It is a statement about data under an assumed scenario, not a statement about how likely that scenario is, and the distinction is where nearly every popular misreading begins.

The conventional threshold below which results are called statistically significant is a convention, adopted for convenience and hardened by habit into a gatekeeping rule. Nothing in nature changes on either side of it. A result marginally above the threshold and one marginally below it are, in evidential terms, nearly identical, yet one is often reported as a finding and the other as no effect. Treating a continuous measure of surprise as a binary verdict discards information and creates a false sense of decisiveness.

Four misreadings worth unlearning

The first misreading is that the p-value gives the probability the hypothesis is false. It does not; it is computed by assuming no effect and asking about the data. The second is that a non-significant result proves there is no effect. Absence of evidence is not evidence of absence, particularly in a small study that could never have detected a modest effect. Reporting that a study found no link, without noting whether it had any realistic chance of finding one, is among the most common errors in science journalism.

The third is that significance indicates importance. It does not; it indicates that the result is unlikely under chance alone, which in a very large dataset can be true of effects far too small to matter. The fourth is treating one significant result as settled. Where researchers test many outcomes, subgroups or time points, some will cross the threshold through chance alone. Studies that pre-specify a single primary outcome, and that state how they accounted for multiple comparisons, are markedly more trustworthy.

Confidence intervals carry more information

A confidence interval expresses the range of values compatible with the data, given the model used. It is more informative than a p-value because it shows both the estimated size of an effect and the precision of that estimate. A narrow interval sitting well away from no effect indicates something reasonably well pinned down. A wide interval that spans from meaningful harm through no effect to meaningful benefit indicates that the study, whatever its headline, has not resolved the question.

The habit worth building is to look at both ends of the interval and ask what each would mean in practice. If the optimistic end represents a benefit that would change how a condition is managed, and the pessimistic end represents no benefit at all, then the honest summary is that the effect might be worthwhile and might be nothing. Reporting that mentions only the central estimate, as almost all reporting does, systematically conveys more certainty than the underlying analysis supports.

Sample size, power, and studies that could not have succeeded

Statistical power is the probability that a study will detect an effect of a given size if that effect genuinely exists. Power rises with sample size, with the size of the effect being sought, and with the precision of measurement. An underpowered study is not merely inconclusive; it is actively misleading in a specific way, because among small studies only those that happen to land on an exaggerated estimate will cross the significance threshold and reach publication.

This produces a counterintuitive consequence. A significant result from a small study, far from being impressive, should often be regarded with more suspicion than a modest result from a large one, because the small study could only have reached significance by overstating the effect. When a striking early finding is followed by larger studies showing something smaller or nothing at all, this mechanism is usually at work rather than any misconduct. Sample size belongs in the second sentence of any responsible report.

Significant but trivial, and the reverse

Statistical significance and practical importance are separate judgements, and only the first is produced by the analysis. With enough participants, a difference of no consequence to anyone can be measured precisely enough to be significant. A shift in a blood marker that leaves no one feeling different, a change in test scores too small to affect a child's schooling, a difference in reaction time invisible in ordinary life: all can be statistically real and practically irrelevant.

The reverse also occurs. A study may find a genuinely meaningful effect but lack the numbers to demonstrate it convincingly, producing a wide interval and a non-significant result that is reported as a dead end. Deciding whether a difference matters requires knowing the units, the baseline, and something about what patients or the public would notice. That judgement belongs to clinicians, policymakers and readers rather than to the statistical test, which is why any report that mentions significance without magnitude has told you almost nothing.

Relative risk versus absolute risk

This single distinction explains more misleading health headlines than any other. Relative risk describes how much a risk changes in proportion to what it was; absolute risk describes how common the outcome actually is. A doubling sounds alarming, but doubling a very rare risk still leaves a very rare risk. The relative figure is the larger and more dramatic number, so it is the one that tends to survive into the headline, while the absolute figure that would let a reader judge relevance is dropped.

The remedy is to look for baseline risk. What proportion of people like me experience this outcome in the first place, and what does the reported change do to that proportion? Where a report offers only percentages of change without any statement of how common the condition is, the reader has been given a number they cannot use. Framing matters too: the same result can be presented as a large proportional reduction or a small absolute one, and both can be technically accurate.

Why nutrition headlines keep reversing

Nutrition is unusually difficult to study, which is why its findings appear to contradict each other so often. Most evidence comes from observational research in which people report their own diets, typically through questionnaires asking them to recall usual consumption. Such recall is imprecise and systematically biased, since people under-report what they consider unhealthy. Randomising thousands of people to eat differently for the decades required to observe chronic disease is neither practical nor, in most cases, ethical.

Diet is also thoroughly entangled with everything else. People who eat more of one food differ in income, education, physical activity, alcohol, smoking and healthcare access. Adjustment handles what was measured, imperfectly. Individual nutrients cannot easily be separated from the whole dietary pattern or from the food that a nutrient displaces. Reading such coverage well means noting whether the study observed or intervened, whether diet was self-reported, and whether the finding concerns a disease outcome or a laboratory marker standing in for one.

Reading epidemic curves responsibly

During an outbreak, case counts plotted over time become the most widely watched statistic and one of the most easily misread. A case count is not a measure of infections; it is a measure of infections that were detected, which depends on how much testing is being done, who is eligible, and how accessible tests are. A rise can reflect expanded testing, and a fall can reflect a public holiday when fewer laboratories reported. Comparing the raw counts of two places with very different testing regimes compares their surveillance systems as much as their epidemics.

Reporting delays create a second trap. The most recent days on any curve are almost always incomplete, because results and death registrations arrive over subsequent days, which makes a genuine plateau look like a decline. Analysts therefore look at the proportion of tests returning positive alongside counts, use rolling averages to smooth weekly reporting rhythms, and watch hospital admissions as a slower but more stable indicator. Treating the last point on a curve as a trend is the single commonest error.

Placebo effects: real, but not what people assume

The placebo effect is frequently described as the mind healing the body, which overstates it considerably. Much of what is observed when an inactive treatment appears to work is not caused by the placebo at all. Many conditions fluctuate, and people typically seek help when symptoms are at their worst, so improvement towards their usual state would have occurred anyway. Participants may also report more favourably to a researcher who has been kind and attentive, and may change other behaviour on entering a study.

A genuine placebo response nonetheless exists, most clearly for outcomes mediated by perception and expectation, such as pain, nausea and fatigue. Expectation, ritual and the clinical relationship can measurably alter how symptoms are experienced. What placebo does not do is shrink tumours, clear infections or repair fractures. This is precisely why controlled trials compare against a placebo rather than against nothing: the comparison isolates the part of the improvement attributable to the treatment itself.

Screening tests, base rates, and false alarms

Screening applies a test to people without symptoms in the hope of catching disease early. Its performance is described by sensitivity, the proportion of those with the disease correctly identified, and specificity, the proportion of those without it correctly cleared. Both can be high while the test still performs poorly in practice, because the question a person actually cares about is different: given a positive result, how likely is it that I have the condition?

The answer depends heavily on how common the condition is in the population being screened. When a disease is rare, the great majority of people tested do not have it, so even a small false-positive rate applied to that large group can generate more false alarms than true detections. This is not a defect in the test but a consequence of arithmetic, and it is why screening programmes are targeted at groups where the condition is common enough for the balance to favour testing.

Two further concepts shape the debate. Overdiagnosis is the detection of disease that would never have caused symptoms in a person's lifetime, which converts healthy people into patients and exposes them to treatment they did not need. Lead-time bias makes survival statistics look better simply because diagnosis happened earlier, without anyone living longer. Decisions about individual screening involve genuine trade-offs and should be discussed with a qualified healthcare professional rather than settled from news coverage.

Holding a headline against its study

A short sequence of checks catches most distortions. Find what kind of study it was and how many people took part. Look for the absolute numbers behind any proportional claim. Note whether the outcome measured is the one that matters or a substitute marker. Check whether this is a first report or part of a consistent body of work. Ask whether the population studied resembles the population the headline addresses.

Then compare the claim in the headline against the claim in the paper's own conclusion, which is frequently far more cautious. Words such as linked, associated and tied are signals that no causal claim was established, whatever the surrounding sentences imply. This is not a counsel of scepticism about research; it is a way of extracting from coverage the information that is genuinely there, and of noticing when the most confident part of a story was added after the science was done.

Sources & References

E

Editorial Team

Editorial

In-house writers and editors producing original explainers, guides, and analysis. Articles cite authoritative public sources where helpful.

Related Articles