Universities are beginning to consider whether generative AI can support marking, moderation and feedback. The governance question is not simply whether a system can produce a plausible mark. It is whether the institution can demonstrate that the judgement remains valid, fair, explainable, reviewable and owned by qualified people.
SPECS position
AI must not independently determine a student’s mark, grade, progression or award. A qualified academic authority must remain accountable, able to examine the student’s work, understand the basis of the recommendation, challenge the output and change the decision before it affects the student.
Why human accountability requires more than approval
Placing a person at the end of an automated process does not, by itself, create effective oversight. Human review becomes ceremonial when reviewers cannot understand the system’s limits, do not have time to examine the underlying work, lack authority to override the output or are encouraged to accept the recommendation routinely.
Effective accountability requires an identifiable decision owner, sufficient academic expertise, access to the relevant evidence, realistic workload, recorded reasons and a functioning route for moderation, correction and appeal.
The current evidence supports caution
A 2026 study reported by Cardiff University compared AI-generated marks with human marks for 50 undergraduate bioscience essays. The researchers reported substantial differences at individual-essay and criterion level, a tendency to compress marks towards the middle and individual differences of up to 40 marks. They concluded that the tested models were not suitable substitutes for human tutors assigning grades to extended written work.
The QAA’s September 2026 assessment initiative frames the central quality question as whether assessment genuinely shows what a student has learned. QAA is developing sector-owned principles and a framework because inconsistent approaches can weaken confidence in award standards.
UNESCO’s guidance places human agency, pedagogical appropriateness, ethical validation and accountability at the centre of educational uses of generative AI. The NIST AI Risk Management Framework similarly expects organisations to define human-AI roles, document oversight, test systems in conditions similar to deployment and monitor performance over time.
What universities should govern
1. Purpose and decision boundary
State precisely whether AI will assist with administrative preparation, rubric checking, formative feedback, moderation signals or mark recommendations. Identify prohibited uses. Open-ended permission such as “AI may assist marking” is not a defensible control.
2. Academic authority
Assign accountability to a qualified academic role. Define who approves the use case, who makes the grade decision, who reviews anomalies and who can suspend the system. Responsibility cannot be transferred to a model, platform or supplier.
3. Assessment validity
Confirm that the proposed use remains aligned with approved learning outcomes, assessment criteria and disciplinary judgement. A technically consistent output may still measure the wrong construct or reward superficial features that do not represent the intended learning.
4. Context-specific validation
Test the system against qualified human markers using representative student work from the intended discipline, level, assessment format and student population. Examine agreement, error distribution, repeatability, outliers and performance across relevant groups. A vendor claim or a general benchmark cannot replace local validation.
5. Data, fairness and accessibility
Establish lawful and ethical controls for student work, personal data, consent, retention, security and intellectual property. Test for differential impact across language backgrounds, student groups and accessible formats. Material disparities require redesign, stronger safeguards or suspension.
6. Meaningful human review
Reviewers need access to the original student work, approved criteria, the AI output and known limitations. They must have time, competence and authority to disagree. Institutions should monitor how often reviewers change an AI recommendation and why. Near-universal acceptance may indicate automation bias rather than reliable performance.
7. Student transparency and challenge
Students should know whether AI contributes to assessment, what role it plays, what information is processed and how to request human review or appeal. Explanations must be accessible and specific enough to support a meaningful challenge.
8. Continuing assurance
Moderation, external examining and assessment boards should examine AI-influenced decisions explicitly. Monitor mark distributions, anomalies, incidents, complaints and subgroup effects. Changes to a model, prompt, rubric, dataset or policy require controlled review and, where material, revalidation.
The SPECS grading-governance test
SPECS has added a twelve-criterion AI-Assisted Grading Governance Test to its AI Quality Readiness Diagnostic. The test examines governance and accountability, assessment validity, equity and data, human review and student rights, and monitoring and improvement.
The test uses five decision signals:
- Pass: evidence supports controlled use within the approved scope.
- Needs action: a correctable weakness must be closed and verified.
- Pilot only: use must remain bounded, supervised and non-consequential while validation or controls are completed.
- Stop: the use must not be deployed or continued for grading until the material condition is resolved and independently approved.
- Not assessed: the institution does not yet have an evidence-based judgement.
Conditions that should stop deployment
- AI independently assigns or materially determines a consequential grade.
- No qualified person can meaningfully review and override the judgement.
- The use has not been validated in the intended assessment context.
- Student work or personal data would enter an unauthorised system.
- Students are not informed or cannot obtain effective human review.
- Material reliability, fairness or accessibility concerns remain unresolved.
- A model or prompt can change without approval, revalidation or traceable records.
Leadership questions
- Who remains personally accountable for the grade, and what evidence supports that accountability?
- Can the institution show that the system is valid and sufficiently reliable for this assessment and student population?
- Can a reviewer understand, challenge and override the recommendation without relying on the supplier?
- What evidence would cause the institution to pause or stop the use?
- Can a student obtain a clear explanation and a genuinely independent human review?
Conclusion
AI may assist carefully defined parts of assessment practice. It must not dilute academic responsibility or weaken the connection between student learning, evidence and judgement. The defensible institutional position is one in which people retain authority, the system’s contribution is transparent, performance is tested, student rights are protected and every consequential decision can be traced to qualified human judgement.
Sources
- Cardiff University — GenAI cannot accurately mark essays in Higher Education, 24 August 2026.
- QAA — Assuring the standard of UK awards in the age of GenAI, 16 September 2026.
- QAA — State of the Nation: AI, assessment and a sector under pressure, July 2026.
- UNESCO — Guidance for generative AI in education and research, 2023; updated 2026.
- NIST — AI Risk Management Framework Core.
This guidance supports institutional quality assurance and improvement. It is not legal advice and does not replace the requirements of the applicable jurisdiction, accreditation framework or institutional authority.
