A raw score is just a count
A raw score is the number of questions answered correctly, sometimes adjusted by subtracting a fraction for wrong answers. On its own it says nothing about difficulty, so two students who took different test forms cannot be compared by raw score alone.
Because standardized tests are given many times a year, each administration uses a different set of questions, called a form. Some forms are inevitably slightly harder or easier than others, even after careful design.
Equating adjusts for form difficulty
Equating is the statistical process that makes a given scaled score represent the same level of ability regardless of which form was taken. It relies on anchor items, a set of questions repeated across forms, whose performance reveals how a new form compares to earlier ones.
If students on a new form answer the anchor items slightly worse on average than students on a previous form did, statisticians infer the new form is somewhat harder, and adjust the raw-to-scaled conversion so it doesn't unfairly penalize test-takers.
The scaled score is the reported number
Once equating is applied, raw scores are converted onto a fixed scale, such as 200 to 800 for a section or a composite range for the whole test. This is the number that appears on score reports, and it is what test scores are scaled to produce.
The scale itself stays constant across years even as the questions change completely. A scaled score of 650 in one year and 650 the next year are meant to represent the same underlying ability level, which raw scores alone could never guarantee.
A percentile is a rank, not a score
A percentile tells a test-taker what share of a reference group scored at or below their level. A percentile of 75 means the test-taker performed as well as or better than about 75 percent of that comparison group, not that they answered 75 percent of questions correctly.
Percentiles are often confused with percentage-correct, but the two measure completely different things. Percentage-correct is about the test content; percentile is about how a score compares to other people who took the test.
The reference group defines the percentile
A percentile is only meaningful relative to a specific comparison group, usually all test-takers within a defined recent period, such as the past three years. Change the reference group and the same scaled score can produce a different percentile.
This is why score reports specify the norm group and the date range used. A percentile calculated against one year's test-takers is not automatically comparable to a percentile calculated against a different year's group.
Item response theory refines the model
Modern large-scale tests often use item response theory rather than simple raw-score conversion. Each question is assigned a difficulty and discrimination value based on how test-takers of known ability historically performed on it.
Under this model, two test-takers who answer the same number of questions correctly can receive different scaled scores if one answered harder questions correctly and the other answered easier ones, because the model weighs question difficulty directly.
Scaling protects year-to-year fairness
Without scaling, a test-taker's reported score would depend partly on luck, whether their particular form happened to be easier or harder than average. Scaling exists specifically to remove that variable from the outcome.
Test organizations treat this as a core fairness requirement, not an optional refinement. A university comparing applicants from different test dates relies on scaling to make those comparisons valid at all.
Concordance links different test scales
Because the SAT and ACT use different scales and different scoring methods, organizations publish concordance tables that estimate an equivalent score on one test given a score on the other, based on studies of students who took both.
Concordance is an approximation, not an exact conversion, because the two tests measure overlapping but not identical skills. Admissions offices treat concordance tables as a rough guide rather than a precise equation.
Guessing penalties change what a raw score means
Some tests historically subtracted a fraction of a point for each wrong answer to discourage random guessing, while others award points only for correct answers with no penalty. The presence or absence of a penalty changes how raw scores should be interpreted.
Many major tests have moved away from guessing penalties in recent years, reasoning that they disadvantaged cautious test-takers more than they deterred genuine guessing. This shift required recalibrating the raw-to-scaled conversion tables.
Section scores and composite scores differ
Many standardized tests report a scaled score for each individual section, such as reading or math, as well as a composite that combines them. The composite is not always a simple sum; it may apply its own rounding or weighting rules.
Understanding the difference matters for interpreting a score report accurately. A strong composite can mask a weak individual section, and some programs weigh specific section scores more heavily than the composite in admissions decisions.
Score reports often include a range, not a point
Because any test contains measurement error, many testing organizations report a confidence band, sometimes called a standard error of measurement, alongside the single scaled score. This acknowledges that a slightly different set of questions might have produced a slightly different result.
A score of 620 might carry an implicit range of roughly 600 to 640, meaning a retake without any change in actual ability could plausibly land anywhere in that band purely due to measurement noise.
Superscoring changes how scores get used, not scaled
Some universities practice superscoring, combining a student's best section scores across multiple test dates into one composite. This is an admissions policy choice, separate from how the testing organization scales any individual sitting.
Superscoring does not alter the underlying scaling process. It simply changes which already-scaled numbers get combined and how, and policies vary widely between institutions, so applicants should check each school's stated approach.
Scaled scores are not IQ-style ability estimates
A scaled test score reflects performance on a specific set of content at a specific point in time, prepared for through study and practice. It is not designed to be a fixed measure of general intelligence, and testing organizations explicitly avoid framing it that way.
This distinction matters because a scaled score can and does change with additional preparation, unlike traits that are meant to be more stable. Score changes across retakes are expected and routine.
Equating requires large, stable data sets
Reliable equating depends on testing large numbers of people and tracking how anchor items perform across many administrations. Smaller or newer standardized tests may have less robust equating simply because they lack years of accumulated data.
This is one reason long-established tests are often trusted more heavily by institutions than newly introduced ones, independent of content quality. Statistical stability itself takes time to build.
A percentile can shift even if raw performance doesn't
Because percentiles depend on the reference group, a test-taker who performs identically to a previous year's version of themselves could see their percentile move if the overall pool of test-takers shifted in ability that year.
This is distinct from the scaled score itself moving, since the scaling process is designed to keep the scaled score stable across forms. Percentile and scaled score can therefore tell slightly different stories about the same performance.
Score scales are set once and rarely changed
When a testing organization redesigns a test, it typically resets the scale and runs a new equating study, because the old scale no longer applies to fundamentally different content. This is why major test redesigns come with explicit notices about score comparability.
Outside of a full redesign, the numeric scale itself, such as the familiar 200-to-800 range, tends to stay fixed for decades, even as the underlying equating adjustments behind each administration continue to shift subtly year to year.
Practice tests use their own scaling, not always official
Official practice tests published by the test-maker generally use validated scaling formulas that mirror the real exam. Third-party practice tests from other companies may use approximate or estimated scaling that does not perfectly match the actual test.
Students relying on unofficial practice materials for score prediction should treat the resulting scaled score as a rough estimate rather than a guaranteed predictor of actual test-day performance.
Scaling doesn't correct for access or preparation gaps
Equating and scaling ensure that a given scaled score means the same thing regardless of test form difficulty. They do nothing to correct for unequal access to test preparation, tutoring, or school resources, which remain separate and well-documented sources of score variation.
This distinction is often lost in public debate about standardized testing. Scaling is a narrow statistical fix for form difficulty, not a broader claim that the testing process is equally fair to every test-taker's circumstances.
What actually matters for test-takers
A scaled score, not a raw score, is what admissions offices and score reports rely on, precisely because it accounts for which form was taken. Treating a percentile as if it were a percentage-correct is the most common misreading of a score report.
Understanding that scaling exists to remove form-difficulty as a variable helps explain why comparing scores across different test dates is considered valid, and why the underlying raw count of correct answers is rarely published or discussed.
Adaptive tests scale as the test proceeds
Computer-adaptive tests select each subsequent question based on whether the previous one was answered correctly, narrowing in on a test-taker's ability level in real time. The final scaled score comes from the difficulty level reached, not just a raw count of correct answers.
Because two test-takers on an adaptive test rarely see the same set of questions at all, equating in this format relies entirely on the pre-calibrated difficulty of each item in the question bank rather than comparing identical forms.
Accommodated testing does not change the scale
Test-takers with approved accommodations, such as extended time, use the same scaled-score range as everyone else. The accommodation changes testing conditions, not the scoring scale or equating method applied to the resulting raw score.
This is a deliberate design choice so that scores remain comparable across all test-takers regardless of accommodation status, and score reports typically do not flag whether an accommodation was used.
Score validity periods are a policy, not a scaling fact
Universities often state that a test score is valid for admissions purposes for a set number of years, commonly around five. This is an institutional policy decision about relevance, separate from the scaling process, which does not expire.
A scaled score from a decade ago still meant the same thing statistically when it was issued; institutions simply choose not to rely on old scores because underlying skills, like math fluency, may have faded since.
Subscores get less rigorous scaling than main scores
Many tests report subscores for narrower skill areas alongside the main scaled score, but subscores are typically based on far fewer questions, making them statistically less reliable and more prone to swinging between attempts.
Test-makers often caution against over-interpreting a single subscore for this reason, recommending it be read as a general indicator rather than a precise measurement the way the main scaled score is treated.
Cut scores work differently from scaled admissions scores
Pass or fail licensing exams, such as professional certifications, use a cut score, a single threshold set through a formal standard-setting study rather than a continuous scale meant for ranking test-takers against each other.
A cut score answers the question of whether a test-taker meets a minimum competency bar, while a scaled admissions score answers where a test-taker falls along a continuous range, a fundamentally different statistical purpose.
Digital and paper versions require cross-mode equating
When a test moves from paper to a digital format, testing organizations run additional equating studies specifically to confirm that a scaled score on the digital version means the same thing as the equivalent score on paper.
This cross-mode equating is necessary because format differences, such as screen reading versus print reading, can subtly affect performance independent of actual ability, and that effect has to be measured and adjusted for separately.
Retaking a test tends to raise scores modestly
Research on repeat test-takers consistently finds a modest average score increase on a second attempt, generally attributed to familiarity with the test format and reduced anxiety rather than a real jump in underlying ability.
Scaling does not attempt to correct for this practice effect, since it is a genuine, if partial, improvement in test-day performance rather than a form-difficulty artifact that equating is designed to remove.
Score choice policies control what a university sees
Some testing programs let test-takers choose which sitting's scores to send to a university, a policy called score choice. This governs disclosure, not the scaling of any individual sitting, which is calculated identically regardless of whether the score is ultimately sent.
Not every university accepts score choice; some require all sittings to be reported. This is another admissions policy layered on top of an already-completed scaling process, unrelated to how any single score was derived.
International test-takers add another equating layer
When a standardized test is administered in many countries, testing organizations must also verify that translated or internationally normed versions produce scaled scores comparable to the original language version, a separate equating challenge from domestic form-to-form differences.
This cross-national equating is generally harder to validate than domestic equating, because cultural and educational-system differences can affect item performance independent of the underlying skill being measured.
Score delays often reflect equating, not grading time
New test forms sometimes take longer to score than established ones because the equating study for a brand-new form requires collecting and analyzing anchor-item data before a final scaled score can be confidently released.
This is different from delays caused simply by manually graded sections, such as essays. A fully multiple-choice test can still see a delayed release if it introduced new content that needs fresh equating.
Score reports separate scaled scores from percentile ranks visually
Most official score reports display the scaled score prominently and list the percentile separately, often in smaller text or a secondary table, precisely because the two numbers answer different questions and shouldn't be read as interchangeable.
Reading both figures together gives a fuller picture: the scaled score shows where a test-taker landed on the fixed measurement scale, while the percentile shows how that placement compares to recent peers.
Not all countries use the same scaling philosophy
Some national exam systems, particularly those tied to fixed university placement quotas, calculate scores in a way closer to a straight rank or percentile from the start, rather than reporting an intermediate scaled score at all.
This reflects a different underlying purpose: allocating a fixed number of university seats based on relative standing, rather than certifying an absolute ability level that could theoretically be met by any number of test-takers.
Understanding scaling helps interpret a low or high score correctly
A single low scaled score reflects performance on that specific form under measurement error, not a fixed verdict on ability, since retakes routinely show movement. A single high score similarly carries a margin of error rather than being an exact, immutable number.
Reading a score report with this in mind, alongside the distinction between scaled score and percentile, gives a more accurate picture than treating either figure as a precise, standalone judgment of a test-taker's capability.
Sources
- College Board: understanding SAT scores β explains scaled scores and the official scoring process for the SAT
- ACT: understanding your scores β describes how ACT raw scores convert to the 1-36 scale via equating
- Wikipedia: item response theory β background on the statistical model used to scale modern standardized tests
- ETS: how tests are developed and scored β explains equating and fairness principles used across large-scale testing programs
FAQ
What is the difference between a raw score and a scaled score?
A raw score is simply the count of correct answers. A scaled score converts that count onto a standardized range after adjusting for the difficulty of the specific test form, so it can be compared fairly across different test dates.
Does a percentile mean the percentage of questions answered correctly?
No. A percentile compares a test-taker's score to a reference group of other test-takers. A percentile of 80 means the score matched or beat about 80 percent of that group, regardless of the actual percentage of questions answered correctly.
Why do standardized tests need equating at all?
Because each test administration uses a different form with slightly different questions, and forms are never perfectly equal in difficulty. Equating adjusts the scoring so a given scaled score reflects the same ability level no matter which form was taken.
Can a scaled score change if I retake the exact same test content?
Test-takers never see identical content on a retake; a new form is always used. Small score fluctuations between attempts are normal and largely reflect measurement error, not necessarily a real change in ability.
Is item response theory the same as a simple curve?
No. A simple curve typically shifts a whole distribution of raw scores by a flat amount. Item response theory weighs each individual question by its measured difficulty and discrimination, producing a more precise, question-by-question adjustment.
Why can two students with the same raw score get different scaled scores?
If they took different test forms, the equating adjustment applied to each form may differ. A form later judged slightly harder will convert the same raw score into a somewhat higher scaled score than an easier form would.
Are percentiles the same across every year for the same test?
Not necessarily. A percentile is calculated against a specific reference group, usually recent test-takers within a set window. If that pool's overall performance shifts, the same scaled score can map to a slightly different percentile.
What is superscoring and does it affect how scores are scaled?
Superscoring is an admissions office practice of combining a student's best section scores from multiple test dates into one composite. It is a policy about which already-scaled numbers to use, not a change to the underlying scaling process itself.
Why do some tests no longer penalize guessing?
Testing organizations found that point penalties for wrong answers discouraged cautious, high-ability test-takers from answering questions they were unsure of, more than it deterred blind guessing, so many major tests dropped the penalty and recalibrated scaling accordingly.
Does scaling account for unequal test preparation access?
No. Scaling only corrects for differences in form difficulty between test administrations. It does not adjust for, and was never designed to adjust for, unequal access to tutoring, practice materials, or school resources.
How is a composite score different from adding section scores together?
Some tests apply their own rounding or combination rules when producing a composite, so it is not always a simple sum of section scores. Score reports typically explain the exact method used for that specific test.
Why do concordance tables between different tests exist?
Because tests like the SAT and ACT use different scales and content emphasis, concordance tables estimate an equivalent score on one test from a score on the other, based on studies of students who took both, though the estimate is approximate.
Is a scaled score a measure of fixed intelligence?
No. A scaled score reflects performance on specific test content at a specific time, shaped heavily by preparation. Testing organizations explicitly avoid presenting it as a fixed measure of general intelligence.
Why might a newer standardized test have less reliable scaling?
Robust equating depends on years of accumulated data from large numbers of test-takers and repeated anchor items. A newly introduced test simply hasn't had time to build that statistical history yet.
Does the numeric scale, like 200 to 800, ever change?
The numeric range itself typically stays fixed for decades. It only changes when a test undergoes a full redesign with substantially different content, at which point the organization resets the scale and runs a new equating study.
About the Author
We reference Wikipedia and other authoritative sources to explain the background and current understanding of this topic.
Loved This Article?
Share it on WhatsApp β Share it on WhatsApp
Get more guides in your inbox β Subscribe to our newsletter for weekly surprising stories from Egypt, Saudi Arabia, Dubai, and beyond.