Search enough data and you will find relationships that mean nothing. Two unrelated quantities can move together by chance, and the more one looks, the more such coincidences appear. Not every pattern is a truth.
The core idea
A spurious correlation is an association that appears in data without any real underlying relationship. It arises from chance, from hidden common causes, or from the sheer number of comparisons made. The pattern is real in the data but empty in the world.
Why it holds
The mechanism is multiplicity. When many possible relationships are examined, some will look strong by chance alone, and cherry-picking these produces convincing nonsense. Enough searching guarantees false positives.
A deeper reading
The turn is that abundance of data worsens the problem. More variables mean vastly more possible pairings, so the harvest of coincidences grows faster than the truths. Big data amplifies spurious signal.
The tension within
Hidden causes also deceive. Two effects of an unseen common cause march together, tempting a direct causal story where none exists. The real driver hides offstage.
Where it fails
The implication demands discipline. One must correct for multiple comparisons, seek mechanisms, and test on new data before trusting a correlation. Skepticism is the price of finding real structure.
The larger point
Spurious correlations are meaningless associations from chance, hidden causes, or excessive searching, and they multiply as data grow. Big data breeds false patterns faster than true ones. Guarding against them requires correction, mechanism, and out-of-sample testing. The principle rewards the patience to state it precisely and the humility to mark its limits. Precision reveals what it truly claims; humility reveals where it quietly fails. Between these two disciplines lies genuine understanding, which is never the possession of a conclusion but the grasp of why the conclusion holds and exactly how far it reaches.