Assessment design
How to Design and Grade Long-Answer Questions
Design prompts and rubrics that separate argument, evidence, subject knowledge, and communication without double-counting.
- By
- MarkingEase Editorial Team
- Published
- Reading time
- 10 minute read
Long-answer questions can reveal how students select evidence, build an argument, and connect ideas. They also create room for inconsistent grading unless the goal, scope, and rubric are explicit. Good design preserves legitimate variation while giving graders common reference points.
Define the assessment goal first
Decide whether the task assesses explanation, evaluation, design, or synthesis. ‘Discuss microservices’ is too open for consistent judgment. A bounded scenario and decision requirement make the intended performance visible.
Structure the rubric around reasoning
An analytic rubric can separate accurate concepts, scenario application, justification, and treatment of trade-offs. Categories must represent different evidence rather than different labels for the same impression.
| Criterion | Marks | Evidence |
|---|---|---|
| Conceptual accuracy | 5 | Accurate characteristics of both options |
| Scenario application | 5 | Connects team size and traffic to the recommendation |
| Justification and trade-offs | 7 | Builds a coherent case and addresses a disadvantage |
| Organisation and clarity | 3 | Presents a traceable argument without rewarding decoration |
Keep content and presentation distinct
Clear writing makes reasoning visible, but polished language should not compensate for incorrect knowledge unless communication is an outcome. Define clarity as a logical sequence and unambiguous references rather than style preference.
- Avoid rewarding one explanation under knowledge and argument.
- Do not infer understanding that is not expressed.
- Allow concise answers full marks when criteria are met.
Award partial credit without fragmenting the essay
Score criterion by criterion, then check whether the total reflects the evidence. A response may have accurate facts but fail to apply them, or reach a sensible recommendation through incomplete reasoning.
| Level | Typical evidence |
|---|---|
| Strong | Recommendation follows from scenario evidence and addresses a trade-off |
| Developing | Relevant reasons but incomplete scenario links |
| Limited | Preference with isolated facts and little reasoning |
| Absent | No relevant justification |
Anticipate defensible positions
In the sample prompt, either architecture can be defended if reasoning respects the scenario. Reward the quality of the case, not agreement with a reference conclusion. Record unacceptable factual claims separately from acceptable differences in judgment.
- List at least two defensible approaches.
- Identify facts that would invalidate an argument.
- Use reference responses as anchors, not templates.
Moderate long answers
Use common anchor scripts at several levels. Graders should cite evidence for each criterion and periodically re-score an anchor to detect drift. If guidance changes, identify earlier scripts requiring review.
- Select anonymised boundary samples.
- Score independently by criterion.
- Discuss differences using response evidence.
- Document accepted interpretations.
- Recheck consistency near grade boundaries.
Practical checklist
- The prompt states a bounded task and context.
- The rubric reflects intended reasoning.
- Criteria do not double-count evidence.
- Presentation marks are justified separately.
- Multiple defensible positions can receive credit.
- Partial-performance boundaries are described.
- Anchor responses support moderation.