iqmetria.
What is IQ

The History of IQ Testing: From Binet to Today

Published July 30, 2026 · 8 min read

Share

The history of IQ testing (IQ stands for intelligence quotient, a score that compares your reasoning ability to the general population) begins in 1904 Paris, not in a psychology lab chasing a grand theory of the mind, but in a government office trying to solve a practical school problem. That practical, almost bureaucratic origin explains a lot about why IQ testing looks the way it does today — and why it's been misused so many times along the way.

Why did IQ testing even start?

The French Ministry of Education had a specific headache: some children were falling badly behind in school, and teachers needed an objective way to identify who needed extra help, rather than relying on a teacher's gut feeling. They asked psychologist Alfred Binet, working with his colleague Théodore Simon, to build a tool for exactly that.

Binet wasn't trying to rank humanity or measure innate genius. He was trying to spot kids who needed support before they fell too far behind. That distinction matters, because much of what came after drifted quite far from Binet's original, fairly modest goal.

The Binet test: the first real intelligence scale

The Binet test — properly the Binet-Simon Scale, published in 1905 — was a series of short tasks arranged by difficulty: following instructions, naming objects, repeating digits, spotting similarities between things. Binet worked out the typical age at which children could complete each task, giving a child a "mental age" that could be compared to their actual age.

A few things about Binet's approach were genuinely ahead of their time:

  • He believed intelligence was influenced by environment and education, not fixed at birth like eye colour.
  • He warned explicitly against using the scale to rank or label children permanently.
  • He saw the test as a screening tool for support, not a verdict on someone's worth or future.

Binet died in 1911, only a few years after publishing his scale, and much of his caution about misuse didn't travel with the test as it spread beyond France.

How the Stanford-Binet changed everything

American psychologist Lewis Terman, at Stanford University, adapted and expanded Binet's work for American children, publishing the Stanford-Binet Intelligence Scale in 1916. This is the version most people picture when they imagine an "old-school" IQ test.

Terman's most consequential change was adopting a formula from German psychologist William Stern: dividing mental age by chronological age and multiplying by 100. That calculation is literally where the term "intelligence quotient" comes from — quotient meaning the result of a division. A ten-year-old performing at a mental age of twelve would score 120.

That formula worked reasonably well for children, whose mental age visibly climbs year by year. It made much less sense for adults, whose mental age doesn't keep climbing past a certain point the same way — a limitation that later scales had to fix entirely (more on that below).

The Stanford-Binet was revised repeatedly over the following decades and is still in use today, now in its fifth edition, though it looks nothing like a simple age-ratio test anymore.

A dark detour: testing at scale, and its misuse

World War I gave IQ testing its first taste of mass deployment. The US Army needed to sort over a million recruits quickly, so psychologists built group tests — the Army Alpha (for literate recruits) and Army Beta (a nonverbal version for those who couldn't read English). This was the first time intelligence testing moved from one-on-one clinical settings to industrial-scale assessment of ordinary people.

It's also where the history of intelligence testing takes its most uncomfortable turn. Some of the same researchers, along with figures like Henry Goddard, used crude test results to support eugenicist arguments — claims about the supposed inferiority of immigrant groups and justification for discriminatory policy. The tests themselves were poorly translated, culturally loaded, and given under stressful conditions to non-native English speakers, yet the results were treated as objective fact.

This history is worth knowing because it's the root of a fair criticism that still gets raised today: that intelligence tests can be culturally biased if they're not built carefully. Modern test design takes this seriously, and it's a large part of why psychologists moved toward tasks that rely less on language and cultural knowledge — the kind of abstract reasoning puzzles you see in many tests now.

Wechsler and the birth of the modern IQ scale

The next major leap came from psychologist David Wechsler, who in 1939 pointed out the obvious problem with the age-ratio formula: it collapsed for adults. He proposed comparing a person's raw score to the scores of other people the same age, rather than to a moving "mental age" target. This is called a deviation IQ, and it's the method every serious test uses today.

Here's the core idea, in plain terms: a test is given to a large, representative sample of people, and their scores form a bell curve (a normal distribution where most people cluster near the middle and fewer people sit at the extremes). The average is fixed at 100, and a standard deviation (a measure of how spread out scores are around that average) is fixed at 15 points. Your score simply tells you where you land on that curve relative to everyone else your age. You can see exactly how that maths plays out in the IQ bell curve and how scores translate into rankings on the IQ percentile guide.

Wechsler's tests — the WAIS (Wechsler Adult Intelligence Scale) for adults and WISC for children — became, and remain, the gold standard used by clinical psychologists. For a side-by-side look at how these compare with other formats, see the breakdown of WAIS, Raven, and Cattell tests.

From Spearman's g to modern theories of intelligence

While the tests themselves evolved, so did the theory behind them. British psychologist Charles Spearman noticed in 1904 — the same year Binet was starting his work — that people who did well on one type of mental task tended to do well on others too. He proposed a general factor, now called g, that underlies performance across different reasoning tasks.

That idea still anchors most modern IQ tests, though the field has added nuance since. Raymond Cattell later split general intelligence into fluid intelligence (raw, on-the-spot reasoning) and crystallized intelligence (accumulated knowledge and skills) — a distinction explained in more depth in fluid vs crystallized intelligence. Today's tests typically combine several question types — verbal, numerical, spatial, and abstract reasoning — precisely because no single question type captures the whole picture of the g factor.

A quick timeline

YearMilestoneWhat changed
1904Spearman proposes gSuggests a general reasoning factor behind mental ability
1905Binet-Simon ScaleFirst practical intelligence test, built for schools
1916Stanford-BinetIntroduces the mental-age ÷ chronological-age formula and the term "IQ"
1917–18Army Alpha/BetaFirst mass group testing; also first major misuse via eugenics
1939Wechsler-BellevueIntroduces deviation IQ, replacing the mental-age formula
1949–2020sWAIS/WISC revisionsRegular updates to norms and content, still the clinical standard

Why has average IQ kept changing over time?

One quirk that trips people up: if IQ is normed against a population, why do older tests seem "easier" than newer ones? This is the Flynn effect — the well-documented rise in raw test scores across generations, which is why tests are periodically re-normed rather than left static. It's a fascinating side note in the history of intelligence testing, and worth reading on its own in the Flynn effect explained.

What hasn't changed since Binet

Despite over a century of revisions, a few things remain constant:

  • Tests still compare an individual against a representative group, not an absolute standard.
  • Reasoning ability, not accumulated facts, is still the core target.
  • No serious psychologist treats a single score as a complete verdict on someone's worth or potential — that was Binet's warning in 1905, and it still holds.

Key takeaways

  • IQ testing started in 1905 with Binet's practical school screening tool, not a grand theory of intelligence.
  • The Stanford-Binet (1916) introduced the mental-age formula that gave us the term "intelligence quotient."
  • Early mass testing was misused to support eugenicist arguments — a genuine dark chapter worth knowing.
  • Wechsler's deviation IQ (1939) replaced the age-ratio formula with the population-comparison method still used today.
  • The theory behind tests — Spearman's g, later refined by Cattell — still shapes how modern tests are built.

None of this history changes what a test can responsibly tell you about yourself today: a snapshot of reasoning strengths, not a medical or clinical diagnosis. iqmetria's own test, like any reputable modern one, is orientative — a genuinely useful tool for self-knowledge, not a final word on your worth or your ceiling.

FAQ

Who invented the first IQ test?+

Alfred Binet, a French psychologist, created the first practical intelligence test in 1905 with his colleague Théodore Simon. It was designed to identify schoolchildren who needed extra academic support, not to rank people's overall worth or potential.

What is the Stanford-Binet test?+

The Stanford-Binet is Lewis Terman's 1916 American adaptation of Binet's original scale. It introduced the mental-age divided by chronological-age formula, which is literally where the term 'intelligence quotient' comes from. It's still updated and used today, now in its fifth edition.

Why did IQ = 100 become the average?+

Modern tests use a deviation IQ, introduced by David Wechsler in 1939, which compares your score to a large representative sample rather than an age-ratio formula. The average of that sample is fixed at 100 by design, with most people scoring within 15 points either side.

Was IQ testing ever misused historically?+

Yes. During and after World War I, US Army group tests were poorly translated and culturally biased, yet results were used by some researchers to support eugenicist arguments about immigrant groups. It's a well-documented, uncomfortable chapter in the history of intelligence testing.

How is a modern IQ test different from Binet's original?+

Modern tests use population-based statistical norms (a bell curve with mean 100 and standard deviation 15) rather than a simple age-ratio formula, and they combine multiple reasoning types — verbal, numerical, spatial, and abstract — instead of one fixed task list.

Share

Related articles

See your own IQ — free, no sign-up

Twenty questions, an instant score, no account required. The full report is an optional one-time unlock — no subscription.

Take the free test