LEVELS OF MEASUREMENT
To understand “Levels of measurement”, you must know the distinction between DISCRETE AND CONTINUOUS DATA. If you are unsure of these terms, please click here
NOT ALL DATA IS CREATED EQUALLY
Not all numerical data are genuinely mathematical. Some numbers represent precise measurements, whereas others are simply assigned to opinions, experiences or categories so that researchers can analyse them statistically.
For example, a room’s temperature might be measured at 20°C. This is an objective measurement because it refers to a physical property that can be independently verified: different observers using the same accurate thermometer should obtain the same result. Height, weight and reaction time are similar because each is measured using a standardised unit.
By contrast, people’s experience of the temperature is subjective. One person may find the room cold, while another finds it comfortable, even though the objective temperature is identical. Psychological research frequently converts subjective experiences into numbers by asking participants to rate feelings, attitudes or preferences. For example, two people might both rate Johnny Depp’s attractiveness as 9 out of 10, but this does not necessarily mean that they find him attractive to precisely the same degree. Each person is applying an individual standard shaped by their own preferences and experiences.
Furthermore, the intervals between ratings are not objectively defined. The difference between 8 and 9 cannot be assumed to be identical to the difference between 9 and 10. The numbers therefore function more like ordered labels, comparable to “moderately attractive”, “quite attractive” and “very attractive”, than precise units of measurement.
Assigning numbers to subjective judgements makes them easier to code and analyse, but it does not automatically transform them into objective measurements. The same problem applies to happiness ratings: two people who both report a happiness score of 7 may be experiencing very different emotional states. This raises an important question about whether opinions can legitimately be treated as mathematical values. Critics argue that psychology’s reliance on such measures weakens its scientific status and contributes to its description as a “soft science”
WHY THE DISTINCTION MATTERS
The distinction between subjective ratings and objective measurements matters because the type of data collected determines which statistical procedures can legitimately be used. Data that merely classifies or ranks responses generally permits more limited mathematical analysis, whereas data based on standardised units supports meaningful calculations of averages, variability and relationships between variables. Researchers must therefore consider what the numbers actually represent before deciding how they should be analysed and how confidently conclusions can be drawn
WHAT ARE THE DIFFERENT LEVELS OF MEASUREMENT?
WHAT ARE THE DIFFERENT LEVELS OF MEASUREMENT?
As discussed above, not all data are strictly mathematical. Some data come from genuine measurements, while other data are what we might call pseudo-data: opinions, preferences or categories that researchers have turned into numbers so that they can analyse them statistically.
However, these different forms of data do not all tell us the same thing. Some data tell researchers only which category a person belongs to, such as whether they agree or disagree with a statement. Other data allow responses to be ranked, such as ranking Spain first, France second, and Italy third according to preference. Data obtained from genuine measurements can provide still more information, such as the precise difference between two people’s reaction times.
These differences matter because researchers cannot perform the same mathematical operations on every type of data. Before choosing how to analyse their results, they must first establish exactly what their data represent and how much mathematical meaning the values contain. The four categories used to make this distinction are called the levels of measurement
There are four main categories of data. or levels of measurement
NOMINAL
ORDINAL
INTERVAL
RATIO
The four levels of measurement were introduced by the psychologist Stanley Smith Stevens in 1946. He named them nominal, ordinal, interval and ratio. These are not four unrelated types of data. They form a hierarchy, beginning with data that provide very little mathematical information and ending with data that can be measured precisely. Each step up the hierarchy adds something new. Nominal data simply separate people or objects into categories. Ordinal data do this but also place the categories in a meaningful order. Interval data add equal, measurable gaps between values. Ratio data have all these properties as well as a true zero, meaning that a score of zero represents a complete absence of what is being measured.
Each level of measurement adds an additional property.
Nominal data categorises information and counts frequencies
Ordinal data adds ranking.
Interval data introduces equal measurable distances.
Ratio data adds a true zero point.
As the level increases, the amount of meaningful mathematical information increases. This allows researchers to apply progressively more powerful statistical analyses and draw more precise conclusions from the data
The level matters because it determines what can sensibly be done with the data. If the numbers are merely labels for categories, it would make no sense to add them together or calculate their average. If they represent genuine measurements with equal units, however, a much wider range of mathematical calculations can be performed. Researchers must therefore identify the level of measurement before deciding how to describe and analyse their results.
To summarise: Understanding this hierarchy is important in research and data analysis because the level of measurement determines:
How data can be organised
What mathematical operations are meaningful
Which statistical tests can be legitimately used
How confidently conclusions can be drawn from the results
To identify whether data are nominal, ordinal, interval or ratio, ask three questions:
Are the gaps/differences between the categories mathematically equal and meaningful? (In other words: Are the intervals objective, consistent, and arithmetic — or are they subjective opinions
Can the values be meaningfully ordered from smallest to largest? (Is there a clear, natural ranking?)
Does zero represent a complete absence of what is being measured
THE FOUR LEVELS OF MEASUREMENT
NOMINAL DATA
Nominal data place people, objects or responses into separate categories. The categories have no natural order, and numbers, if used, act only as labels. For example, students might be asked which streaming platform they use most:
Netflix
YouTube
TikTok
Disney+
The researcher can count how many students select each platform, but cannot place the platforms in a meaningful numerical order. Netflix is not “greater than” TikTok, and the difference between them cannot be measured.
DO PEOPLE STEREOTYPE FEMALES BY THEIR HAIR COLOUR?
AN EXAMPLE OF NOMINAL DATA: HAIR COLOUR AND PERSONALITY STEREOTYPES
The following example shows how nominal data might be collected in a psychological study. Suppose participants are shown several hair-colour categories, such as blonde, brown, black, red, grey, blue and pink. They are then asked to associate each colour with a personality description, such as fiery, intelligent, academic, friendly, feminist or promiscuous. This produces nominal data because both hair colour and the personality descriptions are categories. A participant who associates red hair with “fiery” is not measuring how fiery someone is. The participant is simply connecting one category with another. Neither set of categories can be placed on a meaningful numerical scale. Brown hair is not greater than blonde hair, and red hair does not fall between blonde and black hair. Similarly, “fiery” is not greater than “intelligent”, and there is no measurable distance between “academic” and “friendly”. The categories identify differences in type, not differences in amount.
ANALYSING NOMINAL DATA
Nominal data are analysed by counting how frequently each category occurs. In this example, the researcher could count how many participants associated red hair with “fiery” or blonde hair with “promiscuous”. The mode can then be used to identify the most common response. The mean and median cannot be calculated meaningfully. A mean requires numerical values that can be added together, while a median requires values that can be placed in order. Nominal categories provide neither. Measures of dispersion, such as the range and standard deviation, are also unsuitable because there are no measurable distances between the categories. Nominal data are usually displayed using a bar chart or pie chart. Histograms and line graphs are unsuitable because they imply an ordered or continuous relationship between values
HOW TO RECOGNISE NOMINAL DATA
To determine the level of measurement (nominal, ordinal, interval, or ratio), ask the following two questions:
Are the gaps/differences between the categories mathematically equal and meaningful? (In other words: Are the intervals objective, consistent, and arithmetic — or are they subjective/arbitrary?)
Can the values be meaningfully ordered from smallest to largest? (Is there a clear, natural ranking?)
To determine whether data is nominal, researchers examine the nature of the response categories. The critical feature is that the categories represent separate labels with no meaningful order and no mathematically equal or interpretable gaps between them.
A NOMINAL DATA EXAMPLE WITH CHEESE
Consider the following survey question:
Q: Do you like cheese?
YES
NO
MAYBE
One participant selects “Maybe”, another selects “Yes”, and a third selects “No”. These answers clearly fall into different categories. There is no logical way to arrange them from highest to lowest. Changing the order of the options (for example, listing them as Maybe, No, Yes) would not change the meaning of the responses. “Maybe” is not higher or lower than “Yes”. Furthermore, there is no consistent or mathematically meaningful gap between the categories. The difference between “Yes” and “No” is not equal to, or comparable with, the difference between “No” and “Maybe” in any arithmetic sense. The response choices simply represent different types of answers a person can select.
HOW RESEARCHERS SCORE NOMINAL DATA
Researchers do not “score” nominal data in the numerical sense used for ordinal, interval, or ratio data. Instead, they simply assign each response to its appropriate category and count the frequency of each category. For example, they might record that 45 participants said “Yes”, 30 said “No”, and 25 said “Maybe”. No arithmetic operations (such as addition or averaging) are applied to the categories themselves. This is what defines nominal data. The term “nominal” comes from the Latin word nomen, meaning “name”. Nominal data organises information into categories that function as labels or names. These categories describe different types of responses, but they do not represent quantities and cannot be arranged in a progressive sequence or hierarchy.
FURTHER CLARIFICATION
A fundamental example is the simple dichotomy of “Yes,” “No and Maybe.” These represent three distinct categories, yet there is no meaningful way to rank them from highest to lowest, nor any measurable difference between them. They simply indicate which category has been selected. Nominal data, therefore, represents the most basic level of measurement in statistics. It only allows observations to be grouped into categories and counted, but it does not allow ordering or meaningful mathematical comparison between the categories.
CODING NOMINAL DATA
Sometimes numbers are assigned to nominal categories purely for convenience when organising data. For example, when collecting information about family pets, the categories might be coded as:
1 = Dog 2 = Cat 3 = Rabbit
These numbers do not represent quantities. They serve only as labels. Assigning the number 3 to rabbits does not imply that rabbits are greater than dogs or cats, nor does it create any mathematical difference between the categories. The numbers simply identify different categories and carry no mathematical meaning.
KEY TAKEAWAY
Nominal data classify responses into distinct groups with no inherent order and no mathematically meaningful gaps between them. It identifies type rather than amount, degree, or difference, making it the most basic level of measurement
.
NOMINAL DATA TYPE QUESTIONS:
Please circle any reason from below that encouraged you to take drugs
PEER-PRESSURE STRESS COPIED-SOMEONE-YOU-ADMIRED BOREDOM TO-LOSE-WEIGHT TO-LOOK-COOL OTHER
Do you like cheese? YES NO MAYBE
MEAL PREFERENCE:
EGG-SANDWICH CHICKEN-SOUP GREEK-SALAD
RELIGIOUS PREFERENCE:
1 = BUDDHIST, 2 = MUSLIM, 3 = CHRISTIAN, 4 = JEWISH, 5 = OTHER, 6= ATHEIST, 7 = AGNOSTIC
POLITICAL ORIENTATION:
LEFT-WING, COMMUNIST, TORY, FASCIST, DEMOCRATIC, REPUBLICAN, LIBERTARIAN, GREEN
EXAMPLES OF NOMINAL DATA
ADVERTISING IN LONELY HEART ADS: How often do females versus males highlight looks or status in their advertisements? This involves categorising ads by content focus without implying any hierarchy among the focuses.
RELATIONSHIP STATUS: Relationship status options: single, married, cohabiting, divorced
ATTACHMENT STYLE: Attachment styles: secure, avoidant, ambivalent, disorganised.
PERSONALITY TYPE: Personality types: emotional stability, ambivert, extrovert, introvert, and neurotic,
COLLECTING NOMINAL DATA IN PSYCHOLOGY: Psychologists collect nominal data by categorising information without implying hierarchy or quantitative value. This involves using surveys or questionnaires with predefined options to gather information on topics such as learning styles, demographic details (e.g., gender or ethnicity), and specific behaviours or traits. For example, a survey might ask participants to select their preferred learning style from visual, auditory, or kinesthetic options, each representing a distinct category with no inherent order. This method allows researchers to effectively identify and analyse different groups or characteristics within their studies.
ORDINAL DATA
ORDINAL LEVEL IN BRIEF
DESCRIPTION
Ordinal data place responses in meaningful order but do not provide an objective measure of the differences between them.
For example, participants might rate how much they like cheese on a scale of 1 to 10. A person who selects 10 is placing their response towards the highest end of the scale, while a person who selects 5 is placing theirs nearer the middle. However, this does not prove that the first person likes cheese more than the second. One person may use the scale generously, while another applies a much stricter standard.
The intervals between ratings are also not known to be equal. The difference between 2 and 3 may not represent the same increase in liking as the difference between 8 and 9. The scores therefore show the order in which participants place their opinions on the scale, but they do not measure liking in standardised, directly comparable units
CONCRETE ILLUSTRATION OF ORDINAL DATA USING ATTRACTIVENESS RATINGS
A clear way to understand ordinal data is through the following structured example. Participants are shown a photograph of a well-known actor, such as Sydney Sweeney, and are asked to rate how attractive they think she is on a scale from 1 to 10. Each participant selects a number where 1 represents very unattractive and 10 represents very attractive. This produces ordinal data because the responses can be placed in a clear order, but the distances between the values are not equal or meaningfully measurable.
THE PRESENCE OF ORDER: Ordinal data has direction and ranking. A rating of 8 is clearly higher than 6, and a rating of 10 is higher than 9. This distinguishes ordinal data from nominal data, where no order exists. However, the numbers do not represent fixed units of measurement. The scale only tells us the position of a response relative to others, not the size of the difference between them.
THE PROBLEM WITH INTERVALS If one participant gives a rating of 9 and another gives a rating of 7, we know the first participant judged the actor to be more attractive. What we cannot determine is how much more attractive. The difference between 7 and 9 is not a standardised or consistent gap. The intervals between numbers are not uniform. The jump from 8 to 9 may feel small to one person and large to another. Each participant applies their own subjective internal standard when using the scale. One person’s rating of 8 may reflect very high attractiveness, while another person’s rating of 8 may reflect only moderate attractiveness. Because the spacing is inconsistent, the data cannot be treated as having equal intervals. When a participant assigns a rating, they are not measuring attractiveness quantitatively — they are simply placing their subjective judgement into an ordered category. The numbers act as labels for rank positions rather than true units of measurement.
HOW ORDINAL DATA IS ANALYSED Ordinal data can be analysed using measures that respect order but do not assume equal spacing:
The median is appropriate because it identifies the middle value in a ranked list.
The mode is also suitable, as it shows the most frequently selected rating.
The mean, however, is problematic. Although it is sometimes calculated in practice, it assumes equal intervals between values, which ordinal data lack. Adding ratings and dividing by the number of responses treats the numbers as consistent units, which can give a misleading impression of precision. Measures of dispersion, such as standard deviation, should be used with caution because they rely on meaningful distances between values. Simpler options, such as the range (highest minus lowest rating), are often more appropriate.
APPROPRIATE VISUALISATIONS: Ordinal data is commonly displayed using bar charts to show the frequency of each rating. Line graphs may sometimes be used when the ordered nature of the data needs to be emphasised. In both cases, the graph should reflect ranking without implying equal spacing between categories.
KEY TAKEAWAY Ordinal data provides information about order and relative position. It allows us to say that one response is “more” or “less” than another, but it does not tell us by how much. The scale offers ranking without precise measurement
HOW TO RECOGNISE ORDINAL DATA
To determine whether data is ordinal, researchers ask two important questions:
Can the values be meaningfully ordered from smallest to largest? (Is there a clear, natural ranking?)
Are the gaps or differences between the ranks mathematically equal and consistent, or are they subjective and unequal?
If the answers can be ranked but the distances between ranks are unequal or not precisely measurable, the data are ordinal.
SUBJECTIVITY AND UNEQUAL GAPS
The defining characteristic of ordinal data is that while a clear order exists, the intervals between ranks are imprecise and subjective. Different people interpret the same point on a scale in very different ways.
Consider this question:
How much do you like Johnny Depp? Least — 1 — 2 — 3 — 4 — 5 — 6 — 7 — 8 — 9 — 10 — Most
One participant rates him a 10, while another rates him a 5. The 5-point difference looks precise, but it is not. For the person who gave a 10, the rating might mean casual enjoyment of a few films. For the person who gave a 5, it might reflect strong admiration after watching many movies. The “gap” between 5 and 10 does not represent the same amount of liking for both participants. The scale only shows relative order (higher or lower), not the exact difference.
The same principle applies to any subjective rating, such as:
How much do you like cheese? Least — 1 — 2 — 3 — 4 — 5 — 6 — 7 — 8 — 9 — 10 — Most
A rating of 6 is clearly higher than 3, but the actual difference in preference is subjective. One person’s 6 might mean they enjoy cheese occasionally, while another’s 6 might mean they eat it regularly. The intervals are not equal or objectively measurable.
KEY TAKEAWAY
Ordinal data has a clear ranking, which distinguishes it from nominal data. However, because the gaps between ranks are subjective and unequal, researchers cannot treat the numbers as precise measurements. This is why only certain statistics (such as the median and mode) are appropriate for ordinal data, while the mean and standard deviation are usually avoided
ORDINAL DATA TYPE QUESTIONS:
Q1: Do you like cheese? NOT AT ALL, NOT MUCH, IT’S OK, IT’S QUITE NICE. I LOVE IT
Or number scales
Q2: Do you like cheese? LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
Q3: Why did you start smoking? Please circle the number below.
PEER-PRESSURE : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
STRESS : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
COPIED-SOMEONE-YOU-ADMIRED : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
BOREDOM : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
TO-LOSE-WEIGHT : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
TO-LOOK-COOL : LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
Q4: How much do you like Italian food? : NOT AT ALL, NOT MUCH, IT’S OK, IT’S QUITE NICE. I LOVE IT
Q5: How much do you like Thai food? Please circle a response below: LEAST -1 2 3 4 5 6 7 8 9 10 - MOST
TYPES OF ORDINAL SCALES
Ordinal data organises responses into a ranked order or hierarchy, but the exact distance or difference between ranks is usually unknown and not necessarily equal. Below are the main types of ordinal scales commonly used in psychology and social research.
LIKERT SCALES Likert scales are one of the most widely used ordinal scales in surveys. Respondents indicate their level of agreement, frequency, or intensity on a statement. A typical example is:
“To what extent do you feel stressed?”
Not at all
Slightly
Moderately
Very
Extremely
The responses create a clear order (from low to high stress), but the psychological distance between “Slightly” and “Moderately” may not be the same as between “Moderately” and “Very”.
CLINICAL ASSESSMENT SCALES: In mental health and clinical settings, symptoms are often rated using ordered categories. For example, the severity of depression or anxiety may be classified as:
None
Mild
Moderate
Severe
These scales allow clinicians to rank symptom intensity while acknowledging that the intervals between levels are not precisely equal.
RANKING SCALE: Participants are asked to rank items in order of preference or importance. In mating preference studies, for instance, individuals might rank traits such as “physical attractiveness”, “kindness”, and “financial stability” from most important to least important. The result is purely ordinal — first choice, second choice, third choice — with no information about how much more important one trait is than another.
SEMANTIC DIFFERENTIAL SCALES This scale measures attitudes using pairs of opposing adjectives (bipolar scales). Respondents rate a concept on a continuum between two extremes, for example:
Good ———— Bad Happy ———— Sad Strong ———— Weak
Although numbers may be assigned underneath (e.g., 1 to 7), the data remains ordinal because the psychological distance between points is not guaranteed to be equal.
GUTTMAN SCALES, also known as scalogram analysis, present statements of increasing intensity or difficulty. The key assumption is that if a person agrees with a more extreme statement, they will also agree with all milder statements below it. This creates a strict hierarchical order and allows researchers to measure the intensity of an attitude or belief.
STAPEL SCALES: Stapel scales ask respondents to rate how well a single adjective describes a subject using a numerical range, typically from -5 to +5, with no neutral zero point. For example:
Innovative +5 +4 +3 +2 +1 0 -1 -2 -3 -4 -5
The scale produces ordered data, but the intervals are not equal, and the absence of a true neutral point distinguishes it from many Likert-type formats.
INTERVAL DATA
INTERVAL LEVEL IN BRIEF
Interval data have three defining features. The values can be placed in order, the intervals between them are equal, and the scale has no true zero. Equal intervals mean that the same numerical difference has the same meaning throughout the scale. For example, the difference between temperatures of 10°C and 20°C is the same as the difference between 20°C and 30°C: both represent an increase of 10°C. However, zero on an interval scale does not represent the complete absence of what is being measured. A temperature of 0°C does not mean that there is no temperature. It is simply a point chosen on the Celsius scale. Consequently, 20°C is 10 degrees warmer than 10°C, but it is not twice as hot. Interval data therefore allow meaningful addition and subtraction because differences between values can be calculated. Ratios involving multiplication and division are not meaningful because the scale does not begin at a genuine zero
A CONCRETE ILLUSTRATION OF INTERVAL DATA USING IQ SCORES
IQ scores are interval data. An IQ test contains tasks with answers that are scored according to fixed rules. These produce a raw score, which is then converted into a standardised IQ score. The standardised scale has equal intervals, meaning that a difference of 10 IQ points has the same mathematical value wherever it occurs on the scale. For example, the difference between IQ scores of 90 and 100 is equivalent to the difference between 120 and 130.
This does not mean that every correctly answered question represents one equal unit of intelligence. Different questions measure different abilities and levels of difficulty. The equal intervals apply to the final standardised IQ scores, not to the individual questions or raw scores.
Differences between IQ scores can therefore be compared meaningfully. A person with an IQ of 130 has scored 30 IQ points above someone with an IQ of 100. However, IQ has no true zero, so these scores cannot be compared as ratios. An IQ of 100 does not represent twice as much intelligence as an IQ of 50.
HOW TO RECOGNISE INTERVAL DATA
To determine whether the data is interval, researchers ask two key questions:
Can the values be meaningfully ordered from smallest to largest?
Are the gaps or differences between the values mathematically equal, consistent, and meaningful?
If the answers to both questions are yes and there is no true zero point, then the data are interval data.
INTERVAL DATA EXAMPLE WITH IQ SCORES
Consider the following IQ scores:
Frank: 210
Blue: 180
Dolly: 150
These scores show the key feature of interval data. The gaps between them are equal and consistent: Blue is exactly 30 points higher than Dolly, and Frank is exactly 30 points higher than Blue. Because the intervals are the same everywhere on the scale, the differences between scores are meaningful and can be compared directly.
WHY THE TRUE ZERO MATTERS – AND WHY IT IS “MADE UP” IN IQ
The biggest point of confusion for students is the idea of a true zero. A true zero means the scale starts from a real point where the thing being measured is completely absent. For example, in height (ratio data), zero centimetres means no height at all — the person would not exist. Because the zero is real. Moreover, we can say someone is twice as tall as someone else, and that statement actually makes sense in the real world. IQ is different. The zero on an IQ test is completely made up (arbitrary). The test creators simply chose zero as a convenient starting number. There is no point on the IQ scale where intelligence is truly absent. Even a person in a vegetative state has some basic brain activity, so they retain some intelligence. The IQ scale was designed so that the average person scores around 100, and each point represents an equal step away from that average. The zero point was never meant to represent “zero intelligence.” It is just an artificial starting line.
Why can’t we say “twice as intelligent”?
Because the zero is arbitrary, the numbers on the IQ scale show only relative positions, not absolute levels of intelligence. When Frank scores 210, and Dolly scores 150, we know Frank is 60 points higher than Dolly (two equal 30-point steps). That difference is meaningful. However, saying Frank is “twice as intelligent” as Dolly assumes that intelligence starts at zero and can be multiplied like a real quantity (such as weight or height). Since the zero on the IQ scale is artificial and does not mean “no intelligence,” doubling any score does not double the actual intelligence. The numbers are not measuring raw amounts that can be proportionally scaled.
In simple terms:
You can add and subtract on interval scales → differences work.
You cannot multiply or divide meaningfully → ratios do not work.
CLASSIC EXAMPLES OF INTERVAL DATA
STANDARDISED TESTS SUCH AS IQ TESTS. These are among the clearest and most widely used examples. The scoring system is deliberately designed so that equal numerical differences reflect equal differences in underlying performance. This allows meaningful comparisons between scores. However, the scale's starting point is arbitrary. A score of zero does not represent the absence of intelligence, which is why IQ scores are classified as interval data rather than ratio data.
TEMPERATURE SCALES (CELSIUS OR FAHRENHEIT) Temperature provides one of the cleanest illustrations. The difference between 10°C and 20°C is exactly the same as the difference between 20°C and 30°C — the intervals are equal and meaningful. Yet 0°C does not mean “no temperature” at all (true absence of heat occurs at absolute zero on the Kelvin scale). This is why Celsius and Fahrenheit are interval scales, while Kelvin is a ratio scale.
PSYCHOLOGICAL RATING SCALES: OFTEN TREATED AS INTERVAL
Psychological rating scales are commonly used to measure experiences such as mood, anxiety, pain and satisfaction. A single rating, such as a score from 1 to 10, is strictly ordinal because the responses can be placed in order, but the intervals between them are not known to be equal. The difference between 2 and 3 may not represent the same change as the difference between 8 and 9.
Nevertheless, researchers often treat these scores as approximately interval, particularly when several items measuring the same construct are combined to produce a total score. This assumes that differences between scores are sufficiently consistent to justify calculating means and using parametric tests such as t-tests and ANOVA. This practice is common and may be statistically reasonable, especially with well-designed multi-item scales and suitable data, but it does not mean that the original responses have genuinely equal intervals. Individual Likert items are generally regarded as ordinal, whereas summed multi-item scales are frequently analysed as though they were interval.
NEUROIMAGING MEASURES (E.G. FMRI BOLD SIGNAL CHANGES) Functional MRI signals usually reflect relative changes in brain activation compared with a resting baseline, rather than absolute levels of neural activity. The intervals between signal values allow meaningful comparisons of increases or decreases in activation. However, there is no true zero that represents “no brain function,” so these measures are best classified as interval data in most research contexts.
KEY TAKEAWAY: Interval data allows you to compare differences meaningfully (addition and subtraction work) and to use equal intervals for statistical analysis. However, it does not allow you to make proportional (ratio) statements, because the zero point is arbitrary rather than a true absence of the measured attribute. Recognising this distinction helps researchers choose the right statistical tests and interpret findings correctly. Interval scales sit between purely ordered data (ordinal) and fully proportional data (ratio).
RATIO DATA
RATIO DATA
Ratio data has all the properties of interval data — order and equal, consistent intervals — plus one crucial addition: a true zero point. This true zero represents the complete absence of the quantity being measured. Because of it, ratio data allows meaningful proportional comparisons, such as “twice as much”, “half as many”, or “three times as frequent”.
A clear way to understand ratio data is through direct physiological and behavioural measurements commonly used in psychology, such as reaction time and heart rate. In a typical cognitive experiment, participants may be asked to press a button as quickly as possible when a stimulus appears on a screen. Their reaction time is recorded in milliseconds. Similarly, in studies of stress or arousal, a participant’s heart rate may be measured in beats per minute using a monitor. This produces ratio data because the values are numerical, the intervals between them are equal, and there is a true zero point.
A reaction time of 400 milliseconds represents the same increase from 300 milliseconds as 500 does from 400. The units are consistent, and differences between values are directly comparable. This satisfies the requirement of equal intervals, as in interval data. However, ratio data go further because zero represents the complete absence of the variable being measured. A reaction time of 0 milliseconds would mean no delay between stimulus and response. A heart rate of zero beats per minute would indicate no cardiac activity. While such values are not typically observed in practice, the scale itself has a true zero, which confers additional mathematical properties on the data.
Because of this true zero point, ratios between values are meaningful. It is valid to say that a reaction time of 400 milliseconds is twice as long as 200 milliseconds, or that a heart rate of 120 beats per minute is double that of 60 beats per minute. This type of comparison is not possible with interval data, where zero does not represent an absence.
IQ is an example. An IQ test first produces a raw score from the questions or tasks completed correctly. That raw score is then converted into an IQ score by comparing it with scores obtained from a standardisation group. The scale is constructed to have an average of 100, rather than beginning at zero intelligence. Consequently, obtaining no correct answers would not automatically produce an IQ of zero; it would produce a very low score or, more likely, no valid IQ score because the person’s performance fell below the test’s measurable range.
Therefore, an IQ of 60 represents a particular position on a constructed scale, not 60 units of intelligence measured upwards from none. The score indicates that the person’s test performance was far below the average of 100, but it does not reveal an absolute quantity of intelligence. Similarly, an IQ of 120 indicates performance above the average, rather than 120 measurable units of intelligence. The numerical difference between the two scores can be interpreted as follows: 120 is 60 IQ points higher than 60. However, the ratio cannot be interpreted because neither score represents a quantity measured from a genuine zero point. An IQ of 120 therefore cannot be said to represent twice as much intelligence as an IQ of 60
HOW TO RECOGNISE RATIO DATA
To identify ratio data, researchers ask the same two questions used for interval data, with one additional check:
Can the values be meaningfully ordered from smallest to largest?
Are the gaps or differences between the values mathematically equal and consistent?
Is there a true zero point that represents the complete absence of the variable?
If the answer to all three is yes, the data is a ratio.
THE TRUE ZERO POINT
Ratio data has equal and consistent intervals like interval data, but the presence of a true zero gives it additional mathematical power. This true zero allows meaningful multiplication and division. In contrast, interval data has an arbitrary zero, so ratios are not meaningful.
KEY EXAMPLES OF RATIO DATA
Common examples include height, weight, reaction time, distance, age, income, and frequency counts. In each case, zero has an absolute meaning. Zero height means no height at all, zero reaction time means no delay whatsoever, and zero occurrences of a behaviour means the behaviour did not happen. This allows statements such as “twice as fast”, “half as heavy”, or “three times as many”.
Note on AQA exams: AQA does not require students to distinguish between interval and ratio data in most questions. Both are generally treated as suitable for parametric tests, and the focus is usually on whether the data has equal intervals rather than on the presence of a true zer
HOW TO RECOGNISE RATIO DATA
To determine the level of measurement (nominal, ordinal, interval, or ratio), ask the following two questions:
Are the gaps/differences between the categories mathematically equal and meaningful? (In other words: Are the intervals objective, consistent, and arithmetic — or are they subjective/arbitrary?)
Can the values be meaningfully ordered from smallest to largest? (Is there a clear, natural ranking?)
THE TRUE ZERO POINT
Ratio data has all the features of interval data — equal and consistent intervals between values — plus one crucial addition: a true zero point. This true zero represents the complete absence of the quantity being measured. Because of it, ratio data allows meaningful proportional comparisons (multiplication and division), such as “twice as much” or “half as many.” In contrast, interval data has equal intervals but an arbitrary zero, so ratios are not meaningful.
KEY EXAMPLES OF RATIO DATA
HEIGHT Consider the heights of three individuals:
Ralphie: 5'9" (175 cm)
Frank: 5'11" (180 cm)
Blue: 6'0" (183 cm)
The intervals between these heights are equal and stable. More importantly, a height of 0 cm would mean no height at all. This true zero allows statements such as “Blue is 8 cm taller than Ralphie” and meaningful ratios (e.g., one person being 10% taller than another).
WEIGHT Zero weight means the complete non-existence of mass. This allows direct proportional comparisons — one object can meaningfully be described as twice as heavy as another.
OTHER COMMON EXAMPLES
Reaction time (0 ms = no delay at all)
Number of stress-related behaviours per day (0 = none occurred)
Distance, age, income, or frequency counts
In each case, zero has an absolute meaning, enabling statements like “twice as fast,” “half as many,” or “three times the distance.”
COLLECTING RATIO DATA IN PSYCHOLOGY
Ratio data in psychology is collected when variables have equal intervals and a true zero point that represents the complete absence of the measured quantity. This true zero allows meaningful proportional comparisons such as “twice as many” or “half as fast”.
BEHAVIOURAL OBSERVATIONS: Directly counting the frequency of specific behaviours produces clear ratio data. Examples include the number of smiles, aggressive acts, or verbal responses in an observation session. A count of zero means the behaviour did not occur at all, which enables precise comparisons such as “twice as many instances”.
SELF-REPORT SURVEYS Surveys that ask for countable quantities also yield ratio data. Questions such as “How many hours per week do you exercise?” or “How many cigarettes do you smoke daily?” have a true zero (no hours or no cigarettes). Although self-reports can be biased or subject to memory error, the question structure still yields ratio-level data.
GENETIC AND BIOLOGICAL MEASURES: Direct biological measures are often ratio data. Regional brain volume from structural MRI scans (measured in cm³) or counts of specific genetic markers have a true zero — zero volume, which means no tissue is present, and zero markers means complete absence. These allow meaningful proportional analysis.
NEUROIMAGING AND COGNITIVE TESTS: Structural neuroimaging produces ratio data because brain volume has a genuine zero point. In cognitive tasks, the number of correct responses and reaction time (where 0 ms indicates no delay) are also ratio measures.
EXPERIMENTAL TASKS In conditioning, learning, or performance experiments, researchers often record response rates or frequencies. Zero responses mean no activity occurred, which allows clear proportional statements such as “twice the response rate”.
HOW TO ANALYSE RATIO DATA
Ratio data have equal intervals, so the mean, range and standard deviation can be calculated meaningfully. This is also true of interval data: in both cases, the difference between values is consistent and measurable.
Provided that their other assumptions are met, parametric tests can also be used with interval and ratio data. These include the t-test for comparing two means, ANOVA for comparing three or more means, and Pearson’s r for testing a correlation between two variables. These tests are not generally appropriate for nominal or ordinal data because categories and ranks do not have equal, measurable intervals.
Ratio data may be displayed using histograms to show distributions, scattergraphs to show relationships, and line graphs to show changes over time. What distinguishes ratio data from interval data is the presence of a true zero.
KEY CONSIDERATION WHEN IDENTIFYING LEVELS OF MEASUREMENT
A persistent source of confusion in research methods is the tendency to classify data based on how it looks at the end of the process, rather than on what was actually measured when the participant gave their response. This confusion is amplified in psychology because responses are frequently converted into numbers for scoring and statistical analysis. Once numbers appear, there is a strong but misleading assumption that the data must therefore be interval or ratio. This assumption is incorrect. The presence of numbers does not determine the level of measurement. The level of measurement is determined by the nature of the response the participant was required to provide. To establish this clearly, it is necessary to separate two stages in the research process:
WHAT THE PARTICIPANT DOES WHEN RESPONDING
WHAT THE RESEARCHER DOES WHEN ANALYSING THE DATA
These two stages are often collapsed into one, which is where the misunderstanding begins. At the point of responding, the participant is engaged in a specific type of task. If they are asked to choose between categories with no inherent order, such as yes or no or male or female, then the data are nominal. If they are asked to choose from an ordered set of categories, such as never, rarely, sometimes, often, always, or a numerical scale, such as 1 to 5, representing increasing intensity, then the data are ordinal. If they are asked to produce a direct numerical value, such as their height, reaction time, or hormone level, then the data are ratios. This classification is grounded in what the participant is actually doing. It reflects the structure of the response itself, not how the researcher later manipulates the data.
The complication arises when multiple responses are combined. In many psychological measures, including questionnaires such as the Body Shape Questionnaire, participants respond to a series of items using an ordinal scale. Each individual response remains ordinal. However, researchers then sum these responses to produce a total score. At this stage, the data takes on a different appearance. A participant might now have a score of 59 or 90. This looks numerical and can give the impression of precision, as though the difference between 59 and 60 represents a consistent, measurable unit. However, this impression is constructed by the scoring process, not by the original responses. The participant did not generate a continuous measurement. They selected from a set of ordered categories. The numbers assigned to those categories are labels that allow the researcher to quantify and combine responses, but they do not transform the underlying data into true interval or ratio measurements.
In more advanced statistical and psychometric contexts, it is sometimes argued that when many ordinal items are combined, the resulting total score can approximate an interval scale. This is based on the idea that multiple items may smooth out irregularities in spacing between categories and produce a distribution that behaves as if the intervals were equal. However, this is an assumption used for analytical convenience, not a literal transformation of the data into a true interval scale. This distinction is critical. The level of measurement is not determined by the final numerical score. It is determined by the structure of the responses at the point they are given. The researcher’s decision to add, average, or statistically manipulate those responses does not retroactively change the nature of the original data.
CRITICAL DISTINCTION
The level of measurement depends on what is recorded at the point of data collection.
If the data records are categorised, the data are nominal.
If it records ordered positions, it is ordinal.
If it records equal intervals, it is an interval.
If it includes a true zero, it is a ratio.
Reclassifying or grouping data later does not change its original level of measurement.
AQA APPROACH TO LEVELS OF MEASUREMENT
Within the AQA specification, the emphasis is deliberately placed on clarity and consistency. Students are expected to classify data based on what is being measured at the point of response, not on how the data might later be analysed or interpreted in more advanced statistical frameworks. This means that the correct approach in an exam context is to ignore the final score and focus entirely on the response format presented to the participant.
If a participant is selecting between categories with no inherent order, the data are nominal.
If a participant selects from an ordered scale, the data are ordinal.
If a participant provides a direct numerical measurement with equal units and a true zero, the data are ratio data.
Questionnaires such as the Body Shape Questionnaire are therefore classified as ordinal in AQA answers, even though the responses are often summed to produce a total score. The summation process does not alter the level of measurement that the student is expected to identify.
AQA does not require students to engage with the more complex argument that aggregated ordinal data can be treated as interval under certain statistical assumptions. Introducing that argument in an exam answer risks overcomplication and, more importantly, risks misclassifying the data according to the mark scheme.
The most reliable method is therefore procedural:
Identify what the participant was asked to do.
Identify the structure of the response options.
Classify the data based on that structure.
This approach aligns with how exam questions are constructed and how marks are awarded. It avoids the common error of being misled by numerical scores and ensures that the classification reflects the actual measurement process rather than the appearance of the final data
LEVELS OF MEASUREMENT USED IN PSYCHOLOGICAL RESEARCH
The level of measurement is determined by how the data is collected, not how it is later grouped or interpreted. The same test can produce different types of data depending on what is recorded. Classification after the fact does not change the level of measurement.
NOMINAL DATA – TESTS AND EXAMPLES (CATEGORISATION ONLY)
DSM / ICD DIAGNOSES
Data is collected as categories based on symptom criteria.
A person either meets the criteria for depression, OCD, schizophrenia, etc., or they do not.
No ordering and no quantity is measured.STRANGE SITUATION (ATTACHMENT TYPES)
Behaviour is observed, and infants are placed into categories such as secure or insecure.
The data collected is the type of attachment, not a score of “how much” attachment.MBTI (MYERS-BRIGGS TYPE INDICATOR)
Produces personality types such as INTJ or ENFP.
These are categories of preference, not measurements.EYSENCK PERSONALITY INVENTORY (WHEN USED AS TYPE CLASSIFICATION)
If individuals are grouped into types such as introverts or extraverts, the data are nominal.
The classification is categorical.ZODIAC SIGNS (TEACHING EXAMPLE)
Assigned using date of birth, which involves numerical calculation.
However, the recorded data falls into the categories (Aries, Taurus, etc.).
The categories are distinct and unordered.
This shows that the method of assignment can involve numbers, but the resulting data can still be nominal.SALLY-ANNE TASK (PASS / FAIL)
The data collected is whether the child passed or failed.
This is categorical, not a measure of degree.KEY POINT
Nominal data records the category to which something belongs.
It does not measure how much of something is present.
ORDINAL DATA – TESTS AND EXAMPLES (RANKING WITHOUT EQUAL INTERVALS)
GLOBAL ASSESSMENT OF FUNCTIONING (GAF)
Individuals are placed along a scale from severe impairment to high functioning.
There is an order, but the gaps between scores are not equal.EATING ATTITUDES TEST / FASCISM SCALE (LIKERT FORMAT)
Responses such as strongly agree to strongly disagree.
These can be ranked, but the point differential is inconsistent.BODY SHAPE QUESTIONNAIRE
Produces levels of concern about body image.
Indicates varying degrees of concern, but not precise differences.DRIVING TEST OUTCOMES
Pass, minor faults, major faults.
Ordered in terms of performance, but not equally spaced.GCSE GRADES (WHEN USED AS RANKS)
Higher grades indicate better performance, but differences are not uniform.
INTERVAL DATA – TESTS AND EXAMPLES (EQUAL INTERVALS, NO TRUE ZERO)
IQ TESTS / WAIS
Scores have equal intervals.
The difference between 100 and 110 is the same as between 110 and 120.
There is no true zero, so ratios are not meaningful.MMPI
Produces standardised scores across personality dimensions.
Differences between scores are interpretable as equal.GAD-7
Produces a summed score treated as an interval.
Equal differences between scores are assumed for analysis.A LEVELS / GCSE NUMERICAL SCORES (WHEN TREATED AS SCALED DATA)
Differences between marks are treated as equal, but there is no true zero of ability.
RATIO DATA – TESTS AND EXAMPLES (TRUE ZERO AND FULL MEASUREMENT)
REACTION TIME (E.G. MARSHMALLOW TEST, COGNITIVE TASKS)
Measured in seconds or milliseconds.
Has a true zero and equal intervals.PHYSIOLOGICAL MEASURES (HEART RATE, SKIN CONDUCTANCE)
True zero exists.
Ratios are meaningful.NUMBER OF BEHAVIOURS OR RESPONSES
Count data with a true zero.
Allows full mathematical operations.
VALIDITY AND RELIABILITY OF DIFFERENT LEVELS OF MEASUREMENT
VALIDITY AND RELIABILITY OF DIFFERENT LEVELS OF MEASUREMENT
The level of measurement matters for reliability and validity because some types of data are precise measurements, while others rely on opinions and feelings. The more subjective the data, the more room there is for disagreement and interpretation, which reduces reliability and can weaken validity. As data becomes more objectively measured, both reliability and validity generally increase.
NOMINAL
Nominal data simply place observations into categories.
Reliability can be weak because the classification often depends on human judgment. Two researchers observing the same behaviour might not classify it the same way. For example, one observer might label behaviour as aggressive while another might interpret it as playful or competitive. Because the decision depends on interpretation, the same observation may not always be classified consistently.
Validity can also be limited because nominal categories often simplify complex phenomena. A category such as “has depression” versus “does not have depression” does not capture severity or variation in symptoms.
ORDINAL
Ordinal data introduces ranking, but the distances between ranks are unknown.
Reliability can still be problematic because people interpret ranking scales differently. If students rate their stress from 1 to 10, a rating of 7 may represent very different levels of stress for different individuals.
Validity is also limited because the numerical differences between ranks are not meaningful measurements. The gap between ratings of five and six may not represent the same difference as the gap between eight and nine.
INTERVAL
Interval data contains ordered values with equal distances between them.
Reliability improves because the measurement is standardised. Temperature measured with the same thermometer should produce the same result when measured repeatedly under the same conditions.
Validity is also stronger when the measurement corresponds closely to the phenomenon being measured. Temperature readings reflect actual thermal conditions. However, interval scales lack a true zero, so ratios cannot be interpreted meaningfully.
RATIO
Ratio data contains all the properties of interval data but also includes a true zero point.
Reliability is typically very high because the measurements are objective and replicable. Measuring someone’s height or reaction time will produce consistent results across different observers using the same instruments.
Validity is also strongest because the measurement directly reflects the quantity being studied and allows meaningful comparisons, such as one value being twice as large as another.
As the level of measurement increases from nominal to ratio, the precision of the data increases. This reduces reliance on interpretation and increases both the reliability and the validity of the measurement.
Make it stand out
Whatever it is, the way you tell your story online can make all the difference.
