Are IQ Tests Biased? What the Word Means in Psychometrics
Two different questions get called bias, and the research answers only one of them. Keeping them apart is most of the work.
Ask whether IQ tests are biased and you will get confident answers in both directions, often citing the same research. That is not because anyone is lying. It is because “bias” means something specific and narrow in psychometrics and something much broader in ordinary speech, and the studies address the narrow version.
Untangling the two gives you a genuinely useful answer to both, so that is what this page does.
The two things called bias
In everyday use, calling a test biased usually means something like it is unfair, it favours some people over others for reasons that have nothing to do with ability. That is a reasonable thing to mean and a hard thing to measure.
Psychometrics uses the word for two specific properties that can be measured:
- Predictive bias. Does the test systematically over- or under-predict some outcome, such as school grades or job performance, for one group compared with another? If two people from different groups get the same score, does the same outcome tend to follow?
- Measurement bias. Does a particular question behave differently for people of equal underlying ability who belong to different groups? This is assessed as differential item functioning.
Both are testable claims with established methods behind them. Neither is the same as the everyday question, and a test can pass both while still producing outcomes people reasonably call unfair. That is not a gotcha, it is just what the terms mean.
What the research finds on predictive bias
This is the question that has been studied most heavily, largely because it is the one that matters legally for hiring and admissions.
The general finding, across a large literature, is that well-constructed ability tests predict outcomes about equally well across English-speaking groups within a country. A given score tends to be followed by similar grades or job performance regardless of which group the test-taker belongs to. Arthur Jensen’s extensive work on this in the 1970s and 1980s reached that conclusion, and the broad position has held up in subsequent work.
Where predictive bias has been detected, it has usually taken a form that surprises people: over-prediction for minority group members rather than under-prediction. In technical terms, studies tend to find intercept differences rather than slope differences, meaning the test predicts slightly better outcomes than actually occur for some groups. That is the opposite direction from what the everyday accusation implies.
The standard consensus statement on this and related questions is Intelligence: Knowns and Unknowns, produced by an American Psychological Association task force chaired by Ulric Neisser and published in 1996. It was commissioned specifically to establish what the field could and could not agree on, and it remains the most cited attempt to draw that line.
How individual questions are screened
Differential item functioning, usually shortened to DIF, is the tool for catching problems at the level of single questions rather than whole tests.
The logic is straightforward. Take people who score the same overall, so their underlying ability is comparable. If members of one group are systematically more likely to get a particular item right than equally able members of another group, something about that item is doing work other than measuring ability.
Modern test development screens for this routinely. Items flagged by DIF analysis are reviewed and often removed before a test is published, which is one of the more meaningful differences between a clinical instrument and a quiz someone assembled over a weekend.
Worth being precise about what a DIF flag means, though: detecting it identifies an item for review rather than proving bias. There can be legitimate reasons why an item behaves differently across groups, and judgement is required. It is a screening signal, not a verdict.
The limit nobody should skip
Here is where a lot of confident writing on this subject stops, and it should not.
Showing that a test predicts an outcome equally well across groups does not establish that the test is fair in the ordinary sense. It establishes that the test relates to that outcome consistently. If the outcome itself is shaped by unequal schooling, unequal opportunity or unequal treatment, then a test that predicts it accurately is faithfully tracking a system that is already uneven.
This is not a critic’s objection smuggled in at the end. Jensen, who argued as strongly as anyone that ability tests were not predictively biased, stated the point directly: a culturally biased test can still show good predictive validity for a culturally biased criterion. The statistical finding and the fairness question are different, and the first does not settle the second.
So the accurate summary is narrower than either camp tends to offer. Well-built tests are not, in general, predictively biased in the technical sense. Whether the whole arrangement is fair is a question about criteria, access and consequences that psychometrics is not equipped to answer.
Curious about your own score? Our 50-question assessment covers verbal, numerical, spatial and memory reasoning. Like any online test it is an indicative estimate rather than a clinical assessment, for reasons set out below.
Cultural loading is real and different
Cultural loading is not the same as statistical bias, and conflating them causes half the confusion in this debate.
An item is culturally loaded when answering it draws on knowledge that is unevenly distributed for reasons unrelated to reasoning ability. Vocabulary items are the obvious case: knowing a word requires having encountered it. General knowledge items are another. So, more subtly, is familiarity with the conventions of multiple-choice testing, working at speed under observation, and the tacit understanding that you should guess rather than leave blanks.
This is why the field distinguishes culture-loaded from culture-reduced tests. Matrix reasoning tasks, which ask you to complete abstract visual patterns, are culture-reduced because they require no specific vocabulary or knowledge. They are not culture-free, and tests marketed as culture-fair are more accurately described as culture-reduced. Familiarity with abstract puzzle formats is itself something schooling provides unevenly.
The unfairness that actually bites
In practice, the things most likely to make a particular score misleading for a particular person are not subtle statistical properties. They are mundane and well understood:
- Testing in a second language. The largest single issue. Verbal subtests measure language proficiency alongside reasoning when the test-taker is not a native speaker, and the resulting composite understates ability. Qualified assessors are expected to account for this, often by relying on non-verbal measures or interpreting verbal indices with explicit caution.
- Unfamiliarity with testing. Someone who has sat standardised tests throughout their schooling has an advantage over someone who has not, independent of ability.
- Testing conditions. Illness, exhaustion, stress, pain and medication all depress performance. So does an examiner the test-taker finds intimidating or unfamiliar.
- Uneven cognitive profiles. A composite score averages several quite different abilities, which can misrepresent someone with a specific learning difficulty. This is covered in our guide to an IQ of 127.
None of these makes the test biased in the technical sense. All of them can make a specific score wrong about a specific person, which is what most people are actually asking about.
Where online tests stand
Worth being straight about this, including about our own.
Clinical instruments like the WAIS-5 and WISC-V go through item-level DIF screening, large representative norming samples and published technical manuals that let anyone check the work. Most online tests do none of that. They are typically normed on whoever happened to take them, which is a self-selected group rather than a representative sample, and they publish nothing about how items were chosen.
That is a construction problem rather than a bias problem specifically, but it has the same practical consequence: you cannot tell what the number means. A well-built online test can give a useful indicative estimate. It is not equivalent to a supervised clinical assessment, and anyone claiming otherwise is overselling. More on that distinction in our guide to online test accuracy.
Related reading
Frequently asked questions
Are IQ tests biased?
It depends which question you are asking, because "bias" has a technical meaning in psychometrics that differs from the everyday one. In the technical sense of predictive bias, well-constructed tests have been found to predict outcomes such as school grades and job performance about equally well across English-speaking groups within a country. In the everyday sense of being culturally loaded or unfair in context, there are real and documented problems, particularly around language, test familiarity and testing conditions.
What does test bias actually mean to psychologists?
Two specific things. Predictive bias means a test systematically over- or under-predicts an outcome for one group compared with another. Measurement bias, usually assessed as differential item functioning, means a particular question behaves differently for people of the same underlying ability who belong to different groups. Both are statistical properties that can be tested for, which is what makes them useful terms.
Are IQ tests culturally biased?
Cultural loading is real and is not the same thing as statistical bias. Items that depend on vocabulary, general knowledge or familiarity with particular formats draw on things distributed unevenly across cultures and schooling. Test publishers screen items for differential item functioning and remove those that behave oddly, but no test is genuinely culture-free, and tests described as culture-fair are better understood as culture-reduced.
Do IQ tests predict outcomes equally well for different groups?
That is the specific claim the research has examined most closely, and the general finding is yes for well-constructed tests used within a country and language. Where predictive bias has been found, it has tended to take the form of over-prediction for minority group members rather than under-prediction, which is the opposite direction from what the everyday use of the word implies.
Does predicting outcomes well mean a test is fair?
No, and this is the most important limit on the research. A test can predict a criterion accurately while the criterion itself reflects unequal access and opportunity. Arthur Jensen, who argued strongly that ability tests were not predictively biased, acknowledged directly that a culturally biased test can still show good predictive validity for a culturally biased criterion. Statistical fairness and substantive fairness are different questions.
What is differential item functioning?
A statistical check on individual test questions. If two people have the same overall ability but belong to different groups and have systematically different chances of answering a particular item correctly, that item shows differential item functioning. Modern test development routinely screens for it. Detecting it flags an item for review rather than proving bias, since the cause may be legitimate.
Is taking an IQ test in a second language a problem?
Yes, a serious one, and it is a validity problem rather than a subtle statistical issue. Verbal subtests in particular measure language proficiency alongside reasoning when the test-taker is not a native speaker. Qualified assessors are expected to take this into account, and in many cases to use non-verbal measures or interpret verbal indices with explicit caution.
Are online IQ tests more biased than clinical ones?
They are more likely to be poorly constructed, which is a related but distinct problem. Clinical instruments undergo item-level screening, large representative norming and published technical manuals. Most online tests do none of those things, are normed on whoever happened to take them, and provide no information about how items were selected. That is worth knowing before drawing conclusions from any online result, including ours.
The bottom line
In the technical sense psychometrics uses, well-constructed IQ tests are generally not biased: they predict outcomes about equally well across English-speaking groups within a country, and individual items are screened for differential functioning before publication. That finding is solid and it is narrower than it sounds.
What it does not establish is that the whole arrangement is fair, because a test can predict an unequal criterion faithfully. And it says nothing about the things most likely to make your own score misleading: testing in a second language, unfamiliarity with the format, poor conditions on the day, or a cognitive profile too uneven for a single composite to describe. Those are the questions worth asking about any particular result.
Want an indicative score?
Our 50-question test covers verbal, numerical, spatial and memory reasoning, and reports a percentile with a breakdown by domain.
Start the free testSources
- Neisser, U., et al. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51(2), 77-101. American Psychological Association Task Force report.
- Jensen, A. R. (1980). Bias in Mental Testing. New York: Free Press.
- Jensen, A. R. (1977). An examination of culture bias in the Wonderlic Personnel Test. Intelligence, 1(1), 51-64.
- Reynolds, C. R., & Suzuki, L. A. (2013). Bias in psychological assessment: An empirical review and recommendations. In Handbook of Psychology. Wiley.
- Wechsler, D. (2024). Wechsler Adult Intelligence Scale-Fifth Edition (WAIS-5): Technical and Interpretive Manual. Bloomington, MN: Pearson.
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing.