Flynn Effect Reversal: What the Data Shows About Falling Cognitive Test Scores
Average scores on standardized intelligence tests rose steadily across the industrialized world for most of the twentieth century. The pattern is named after James Flynn, the New Zealand political scientist who documented it in the 1980s by noticing that test publishers kept quietly re-norming their instruments upward. Each new cohort was outscoring the standardization sample of the previous one, so the tests had to be recalibrated to keep the average at 100. Gains ran somewhere near three points per decade, concentrated in abstract reasoning and pattern recognition rather than vocabulary or arithmetic.
That was the story for roughly six decades. It is no longer the story.
Where the reversal was first measured
The clearest evidence does not come from schoolchildren. It comes from military conscription records in Scandinavia, which are unusually good data because they cover nearly the entire male population of a country, use a consistent test over decades, and can be linked to family records.
Norwegian conscript scores peaked with men born around the mid-1970s and declined for cohorts after that. The Norwegian analysis published in 2018 found something that complicated the usual explanations: the decline showed up within families. Younger brothers scored below older brothers. That finding undercuts arguments based on immigration or changing population composition, because the comparison holds parentage constant.
Similar downturns appeared in Danish, Finnish, and British data. The timing matters for anyone building a causal argument. These declines began with cohorts who reached adolescence in the late 1980s and early 1990s, well before smartphones, social media, or one-to-one device programs in classrooms.
What Gen Z data actually shows
The picture for people born between roughly 1997 and 2012 rests on different instruments, and it is worth keeping them separate because they measure different things.
International assessments of fifteen-year-olds recorded their steepest declines on record in the 2022 cycle, with mathematics falling hardest across most participating countries. Pandemic school closures are an obvious candidate, but scores in several countries had been sliding since around 2012, so closures accelerated a trend rather than starting one.
National long-term trend testing in the United States tells a similar story. Thirteen-year-olds posted their lowest reading and mathematics scores in decades, and the lowest-performing students fell furthest, widening the gap between the top and bottom of the distribution.
Adult skills surveys measuring literacy, numeracy, and problem-solving found a growing share of adults at the lowest proficiency levels in many wealthy countries, including among younger respondents who had completed more formal schooling than their parents.
The common thread is not that every measure is down uniformly. It is that the bottom of the distribution is thinning out while the top holds roughly steady, and that this is happening alongside rising educational attainment.
The screen-time argument
The reversal reached a wider audience in January 2026, when neuroscientist and education consultant Jared Cooney Horvath testified before the U.S. Senate Committee on Commerce, Science, and Transportation. He told senators that Gen Z is the first generation in modern history to underperform its predecessors across attention, memory, and general reasoning, and argued that classroom technology is a direct cause rather than a bystander. His position is that the delivery device is irrelevant. Phone, laptop, or school-issued tablet, and whether or not the software carries an educational label, the effect on learning is negative.
The testimony went viral, drawing millions of views and endorsement from prominent tech-sceptical psychologists, and it moved a self-published book into a major-publisher reissue.
What is contested
Journalists and researchers who worked through the underlying evidence reached a more qualified conclusion. There is no smoking-gun dataset showing that education technology caused the recent declines. The strongest claims rest on correlation between two trends that happen to overlap in time, and the overlap is imperfect in an important way: the Scandinavian reversals predate the technology by fifteen to twenty years.
Three further objections come up repeatedly.
The first is measurement. Flynn himself argued that twentieth-century gains reflected a shift in habits of mind toward abstract classification, driven by schooling and cognitively demanding work, rather than a rise in raw intelligence. If that is right, a reversal signals a shift in which mental habits get exercised, not a population getting less capable.
The second is test engagement. Low-stakes assessments depend on students trying. Falling motivation to complete a test that carries no consequences would produce declining scores without any change in underlying ability, and there is evidence that engagement with long assessments has dropped.
The third is confounding. School closures, teacher shortages, changes in curriculum and assessment policy, air quality, sleep duration, and family structure all moved during the same window. Isolating one variable from that tangle requires designs that mostly have not been run.
There is also a conflict-of-interest question that critics raise about advocacy in this space generally, since much of the loudest commentary comes from people who consult, publish, or speak professionally on the topic.
Why the distinction matters
Two claims are travelling together and they are not the same claim.
The descriptive claim is that measured cognitive and academic performance has fallen for recent birth cohorts across several countries and several instruments. That claim is well supported.
The causal claim is that screens and classroom technology are the reason. That claim is plausible, widely believed, and not established.
Policy tends to follow the causal claim. Phone bans in schools, restrictions on one-to-one device programs, and rollbacks of adaptive learning software are all being justified by reference to a mechanism that the evidence supports weakly. Some of those policies may be sensible on other grounds, including attention in the classroom and adolescent mental health. But a policy built on a shaky causal story is fragile, and if scores fail to recover after the devices go away, the entire argument for the intervention collapses along with it.
The Flynn effect took decades to identify and longer to explain. Its reversal is being explained in real time, under political pressure, with data that is not yet up to the job.