Reliability and validity are separate
Reliability is whether a procedure gives consistent results — repeat it the same way and it produces the same findings. Validity is whether it measures what it claims to, so findings are accurate. They are independent: a measure can be reliable but not valid — the same answer every time, but the same wrong answer. So consistent is not the same as correct, and every evaluation should say which of the two it means.
Sampling and designs affect validity
Sampling (11.1.7a) affects how far results generalise: random and stratified samples are more likely representative; opportunity and volunteer are biased. Designs (11.1.7b): repeated measures risks order effects (fix by counterbalancing); independent measures risks participant variables; matched pairs improves on both. Reliability comes from standardised procedures — same instructions, timings, materials — so it can be repeated.
Quantitative vs qualitative trade-off
Quantitative (11.1.7c) is numerical data (scores, counts, times). Scored by a fixed rule, it is high in reliability but weaker on validity — a number can miss what behaviour really means. Qualitative (11.1.7d) is non-numerical data (words, themes). It is high in validity but weaker on reliability — it depends on the researcher's interpretation, so themes may differ. Pattern: quantitative reliable but may lack validity; qualitative valid but may lack reliability.
Drawn from real examiner reports.
Reliability is standardisation, not controls
Reliability comes from standardisation — same instructions, timings and materials, so a repeat gives the same result. Asked how it is made reliable, candidates drift to any term they know: controls, participant variables, sampling or ethics. Give a concrete point — same word list, same viewing time, same instructions. "Keeping things fair" earns nothing.
Standardisation was reported as not understood in June 2022 P2 Q1b, where candidates answered on controls, participant variables, sampling or ethics; only concrete points (same trigrams, viewing time, recall time) were credited.
Qualitative vs quantitative data
This is about the form of the data: qualitative is words, quantitative is numbers. Candidates muddle them, and some think data must be numbers. Decide from the data: scores, counts and times are quantitative; words, descriptions and themes are qualitative. It matters because quantitative tends to be reliable, qualitative valid.
Qualitative and quantitative data are reported as muddled in June 2024 P2 Q5b and Q2c.
Sampling must link to representativeness
Analysing a sampling method's validity means saying why it does or does not give a representative sample, not just naming it. Stratified sampling is often weak: divide the target population into subgroups, then take from each in proportion to size. Weak answers drift into opportunity or random. The mark is in linking method to representativeness.
Stratified sampling drifting into opportunity or random is reported in June 2024 P2 Q2b.
Reliable does not mean correct
A damaging error is treating reliable and valid as the same word. Reliability is only consistency — a reliable measure gives the same result when repeated, which says nothing about whether it is true. A scale reading two kilograms heavy is reliable but not valid: the answer is wrong. So "same result, therefore accurate" is false — consistency is not accuracy.
A bigger sample isn't more representative
A common error is "make the sample bigger so results are more valid". More numbers do not fix representativeness: a biased sample — all volunteers, or one age group — stays unrepresentative when you add more of the same kind. What improves generalisation is the nature of the sample, not its size — a stratified sample reflects each subgroup in proportion.
That "more in a sample makes it more representative" misunderstands generalisability — it is the nature of the sample that matters — is reported in June 2019 P2 Q3c and June 2022 P1 Q5b.
Primary vs secondary is about source
Do not confuse the form of data (qualitative/quantitative) with its source. Primary data are collected first-hand for this study; secondary data were already gathered by someone else. Candidates muddle these, and some assume data must be numbers, quoting figures instead of naming the type. The test is who gathered the data — new is primary, reused is secondary.
Primary and secondary being muddled, with candidates quoting figures instead of naming the type, is reported in June 2019 P2 Q6d and June 2022 P2 Q3c.
Sampling techniques get muddled
Beyond stratified, the techniques are systematically confused. Opportunity (approach whoever is available) is muddled with volunteer (people reply to an advert). Random is often given tautologically as "random", when it means every member of the target population has an equal chance of selection. Learn each by its defining action — a blurred definition earns nothing.
Sampling techniques being systematically muddled — volunteer with opportunity, opportunity left blank — is reported in Nov 2020 P2 Q3b and Nov 2021 P2 Q02b.
Name it, state the threat, say the effect
A bare "this is not valid" earns little. Make each point in three moves: name the idea (reliability or validity), state the threat (order effects, a biased sample), then say the effect on the findings — the "so what". End with a supported conclusion.
Two questions sort reliability from validity
Sort each point before writing it. "Same result if repeated?" = reliability. "Measuring what I claim, and reflecting real life?" = validity. So you don't label an order-effects problem "unreliable" when it really threatens validity.
Take the example from the scenario
When a question says "use an example from the scenario", the example must come from it — a general or invented one earns nothing. Quote the study's own detail: its real sample, items or task. Tie the point to what this study did, not to methods in general.
Reliability — whether a procedure gives consistent results: repeat the study in the same way and you get the same findings.
Validity — whether a procedure measures what it claims to measure, so the findings are accurate and reflect real behaviour.
Standardisation — keeping the instructions, timings and materials identical for every participant, which is the concrete route to reliability.
Full notes, flashcards, Q&A and the topic quiz for every premium subject.
Premium plans are US$8.99/month or US$49.99/year — first month free.
Studying with a parent's blessing? Show them this.