Primary vs secondary: who collected it
Both definitions turn on who collected the data and why.
Primary data — the researcher collects it themselves, first-hand, for their own study.
Secondary data — data already collected by someone else, for a different purpose, that the researcher reuses (a census, a newspaper, an earlier study).
The separator is the source. A vague "primary is your own data" scores poorly — name both who collected it and what for.
Each type has a strength and a weakness
Primary — strength: collected for this study, so it fits the aim and the researcher controls how it is gathered; weakness: slower and more expensive to obtain.
Secondary — strength: quick and cheap because it already exists, and it can cover very large samples or long periods; weakness: collected for a different purpose, so it may not fit the aim and its quality cannot be controlled.
Choose the type that suits the aim.
Source is separate from form
Primary versus secondary is about the source (who collected the data and why). That is separate from whether the data is numerical or in words — the qualitative/quantitative form. The two axes are independent, so all four combinations exist: primary can be quantitative (counting recall) or qualitative (an interview you run); secondary can be quantitative (crime figures) or qualitative (someone else's diaries). Decide the source first, then the form.
Drawn from real examiner reports.
Primary and secondary swapped over
Candidates swap the two, calling secondary data "data you collect yourself" or primary "data from somewhere else". The anchor is first-hand: if the researcher personally gathered the data for this study it is primary; if they reused data someone else had already gathered, it is secondary. Ask: did this researcher run the study that produced the data?
Primary and secondary data being muddled is reported in June 2019 P2 Q6d and again in June 2022 P2 Q3c.
Quoting a figure instead of naming the type
Asked to identify the type of data, candidates copy a number from the scenario — "42 participants" — instead of writing primary or secondary. Data need not be numerical, and a figure is not the name of a data type. "Identify the type" wants the label; only when the question also says justify do you point to the detail of who collected it.
Some candidates thought data must be numerical and quoted figures instead of naming the type in June 2019 P2 Q6d.
Two definitions score only 1 mark
For "describe the difference", writing one definition then the other side by side is treated as a single point and earns 1 of 2 marks. The second mark needs a comparative connective: primary data is collected first-hand, whereas secondary data was collected by someone else for another purpose. The connective makes it a difference, not two definitions.
Answering "describe the difference" with two side-by-side definitions, where connectives such as "whereas" earn the second mark, is reported in June 2019 P1 Q6 and June 2022 P1 Q7.
A vague "your own data" is not enough
"Primary is your own data" and "secondary is other people's data" are too vague to score well. The mark-scheme definition names who collected the data and what it was collected for: primary is gathered first-hand for this study; secondary was collected by someone else for a different purpose. Add the purpose, not just the ownership.
Analysing data does not make it primary
Using or analysing data in your own study does not turn it into primary data. The type is fixed by who originally collected it and why, not by what the current researcher does with it. Census figures a researcher analyses are still secondary, because the census was gathered by someone else for another purpose — the analysis is just what she does with existing data.
Secondary data is not second-rate data
Because primary data is "your own", students assume secondary data is second-rate. It is genuine data with real strengths: it already exists, so it is fast and cheap, and can span far larger samples and longer periods than one researcher could gather first-hand. Neither type is automatically better; which one is depends entirely on the aim of the study.
Source is not the form (numbers vs words)
A frequent muddle is treating "primary" as numbers and "secondary" as words. The source (who collected it) has nothing to do with the form. Primary data can be quantitative (counting recall) or qualitative (an interview); secondary can be quantitative (statistics) or qualitative (someone else's transcripts). Ask who collected it, then whether it is numbers or words.
Name the type, then justify from scenario
Scenario questions reward a two-part answer: first name the type, the AO1 point; then justify it from the scenario — for example "she collected the questionnaires herself" shows primary data. The justifying detail must come from the study, not a pre-learned example.
A character's name is not application
Naming the type and then repeating the character's name does not earn the application mark — the name alone is not enough. Point to what the researcher did: gathered the data first-hand, or reused records someone else had collected. Quote the action, not who performed it.
No AO2 link, no AO3 mark
On a 4-mark strength or weakness, a pre-learned point with no link to the scenario scores zero — AO3 cannot be credited without AO2. Tie each point to this study: why primary data suited this aim, or why secondary was risky here. Application first, evaluation second.
Primary data — data the researcher collects themselves, first-hand, for their own study.
Secondary data — data that has already been collected by someone else, for a different purpose, which the researcher then reuses.
The word that separates them is the source: did this researcher gather the data first-hand, or did they take data that already existed?
Ask one question: did this researcher run the study that produced the data?
Full notes, flashcards, Q&A and the topic quiz for every premium subject.
Premium plans are US$8.99/month or US$49.99/year — first month free.
Studying with a parent's blessing? Show them this.