Testing vs evaluation, against the spec
Testing checks how the product (or prototype) performs against each specification point — the measurable side: capacity, mass, does it work, survives a drop test. Evaluation is the judgement built on those results: for each point you weigh the good AND bad and decide how well it is met, ending in a judgement of fitness for purpose. Work point by point with a reason — for existing products (to learn from) and your own work — not a vague overall opinion.
User feedback — sampling, questionnaires, interviews
Some spec points — comfort, appeal, ease of use — can't be measured on a bench; they need user feedback. Sampling tests a representative group of intended users, not everybody. A questionnaire gives closed questions to many people, quick and easy to chart — it gives breadth. An interview is one-to-one with open questions: slower, fewer people, but it probes the reasons behind an opinion — depth. A user trial is a further method.
Suggest modifications — the five types
Testing and evaluation only matter if they drive improvement. A full evaluation ends by suggesting modifications — specific changes that meet the spec better, in five types: functional (works better), safety (remove a sharp edge, add a guard), aesthetic (look, colour, finish), ergonomic (fits the user better), economic (less cost, material or parts). Justify each by a test result, shown as sketches and notes.
Drawn from real examiner reports.
Describing the product, not evaluating it
The most-repeated evaluate error, in every Paper 1 variant: candidates describe a product ("it has a handle and a lid") when asked to evaluate it. Evaluate weighs the good AND bad against the specification and reaches a justified judgement ("the wide base is stable, but the tall shape wastes material"). Only that scores; saying what a product is or has does not.
Describe-vs-evaluate, a key message every variant (s23 P11–P13 Q(d)).
Testing but stopping before modifications
In coursework AC7 and Paper 1 evaluate parts, candidates test against the specification then stop — reporting pass/fail but never proposing further development. The top band needs a conclusion AND specific modifications (what to change and why), best shown as sketches and notes. "The lid leaked, so add a rubber gasket" has modified; "the lid leaked" only tested.
Testing but no proposals for improvement (s23/w23 P02 AC7).
Opinion, not tested against the spec
Evaluation often is not anchored to the spec or to user evidence: candidates write what they think ("I like it") instead of testing point by point and reporting each result. Coursework research with no relevant user data (interviews, surveys, measurements) stays in the lower band. Fix: take each spec point, say how you tested it, give the result, judge if met.
Un-anchored, opinion-only evaluation stays low (s23/w23 P02 AC2, AC7).
Repeating one evaluation for every idea
Asked to evaluate three ideas (or several spec points), candidates give the same generic judgement for each instead of a distinct, justified verdict per idea. Each idea meets the brief differently, so each needs its own weighing of good and bad against the spec. A copy-paste "this works well" across all three shows no real comparison and caps the marks.
One evaluation repeated for all ideas (s23 P12).
Unjustified assertion with no reason
"It works well", "it is the best idea" and "it looks good" score nothing on their own — they are assertions with no reason tied to a specification point. Every judgement needs the because: why it works, measured against which requirement. "It is stable because the wide base lowers the centre of mass, meeting the tip-test point" earns the mark; "it is good" does not.
Unjustified assertions score nothing (s23 P12 Q2d, P13 Q3d).
Bigger questionnaire ≠ deeper insight
Assuming "more people means better feedback", so a large questionnaire must beat an interview. The two trade breadth against depth. A questionnaire reaches many people and shows how common an opinion is; an interview reaches few but uncovers the reasons behind it. Neither is simply better — use a questionnaire for WHAT users think, interviews for WHY.
Evaluation is just a personal verdict
A testable wrong belief: that evaluating means saying whether you like it or whether it works. Real evaluation is measured against the specification — each point tested, the result reported, good AND bad weighed to judge fitness for purpose. "It works" covers only one point (function) with no reason; a product can work yet fail on cost, safety or appeal.
Structure a top-band evaluate answer
Work against the spec, one point at a time; state how you tested each — a measurement for objective points, a questionnaire/interview for user points; give the result, weighing a good AND a bad with a reason; end each weak point with a named-type modification.
Choose the user method to fit the question
Match the feedback method to what you need. Use a questionnaire with a representative sample to find how common an opinion is (breadth, quick to chart); use interviews to find why users think it (depth). Name the method and say which spec point it tests.
Always name the modification type
When you propose a change, label it — functional, safety, aesthetic, ergonomic or economic — and tie it to the test result that prompted it. Sketch it where you can. Proposals for further development are the top-band evidence, so never stop at "I would improve it".
Full notes, flashcards, Q&A and the topic quiz for every premium subject.
Premium plans are US$8.99/month or US$49.99/year — first month free.
Studying with a parent's blessing? Show them this.