Generative AI has broken a great deal of traditional assessment, and pretending otherwise helps no one. When a student can produce a competent essay, problem set solution or code snippet in seconds, the take-home tasks that dominated assessment for decades no longer reliably evidence learning. Academics worldwide are grappling with the same urgent question: how do we assess genuine learning when students have powerful AI at their fingertips?
The answer is not surveillance, panic or a retreat to nothing but invigilated exams. It is a considered move toward authentic assessment — tasks that measure the deep, transferable capabilities that matter, in ways that are meaningful whether or not AI exists. This article offers practical guidance for academics in any discipline and any country. The challenge is universal, and so are the best responses. For related teaching resources, see our Teaching & Assessment hub.
Why traditional assessment is under pressure
Much traditional assessment was always a proxy. We could not directly observe whether a student could think, analyse or synthesise, so we asked them to write an essay or solve a set of problems and inferred their capability from the product. That inference worked reasonably well as long as producing the product genuinely required the underlying capability. Generative AI has severed that link: the product can now be produced without the capability, so the proxy no longer holds.
This is uncomfortable but also clarifying. It forces us to ask what we actually want students to be able to do, and to assess that directly rather than through a now-unreliable proxy. Many of the tasks AI can complete easily were, in truth, testing low-level recall and reproduction rather than genuine understanding — exactly the things employers and society least need from graduates. The disruption, painful as it is, is an opportunity to make assessment more meaningful, not merely more defended. The wrong response is to double down on tasks AI can do; the right one is to assess the human capabilities that remain valuable.
The principles of authentic assessment
Authentic assessment asks students to apply knowledge to realistic, meaningful tasks that resemble what practitioners actually do, rather than artificial exercises that exist only to be graded. Its principles predate AI but have become newly urgent. Good authentic assessment is contextualised in real-world scenarios, requires judgement and the integration of multiple skills, produces something with genuine purpose, and often involves process as well as product.
Crucially, authentic tasks are naturally more resistant to AI shortcuts — not because they are AI-proof (little is) but because they require the student's own reasoning, context, reflection and defence in ways that generic AI output cannot supply well. Assessing the process of learning (drafts, reflections, decisions, iterations) alongside the final product, and requiring students to connect work to their specific context and experience, makes authentic assessment both more meaningful and more robust. The internationally influential work on assessment for learning, summarised by bodies such as Scotland's QAA Enhancement Themes, underpins these principles. See also our guide to managing the workload this redesign involves.
Redesigning your assessment tasks
Practical redesign starts with auditing your existing assessments against a simple test: could a student get a good grade using AI without genuinely learning? Where the answer is yes, redesign. Effective moves include grounding tasks in a specific local or personal context that AI cannot know, requiring students to apply concepts to their own experience or a live case, and assessing oral defence or live application where students must demonstrate understanding in real time.
Other robust approaches include process portfolios that capture the journey of a project with reflections at each stage; in-class or synchronous components that anchor take-home work to demonstrated capability; and tasks that require students to critique, improve or build upon AI output, which turns the tool into the object of assessment. Vary your methods rather than relying on any single "AI-proof" format — a diverse assessment diet is both more robust and more educationally valuable. Our wider teaching and assessment resources offer more task designs.
Using AI openly and teaching AI literacy
The most forward-looking response is not to ban AI but to bring it into the open and teach students to use it well. AI is now part of the professional world our graduates will enter, and pretending it does not exist serves no one. Increasingly, the goal is not assessment despite AI but assessment that includes appropriate, transparent AI use as a demonstrated skill in itself.
This means designing tasks where students use AI openly, then critically evaluate, correct, extend and take ownership of the output — demonstrating the human judgement that turns raw AI generation into genuine work. Require students to document how they used AI and to reflect on its limitations. Teaching AI literacy — understanding what these tools do well and badly, their biases, and the ethics of their use — is fast becoming a core graduate capability, and assessment is where students learn it most deeply. UNESCO's international guidance on AI in education offers a useful framework for setting expectations.
Maintaining academic integrity
Academic integrity remains essential, but the strategy must shift from detection to design. AI-detection software is unreliable, produces false positives that can wrongly accuse students (disproportionately harming those who write in a second language), and is quickly outpaced by new tools. Building your integrity strategy on detection alone is building on sand. The more durable approach is to design assessment that makes dishonesty difficult and pointless, and to build a genuine culture of integrity.
That means being explicit and consistent about what AI use is and is not permitted for each task, explaining the reasoning so students understand rather than merely comply, and treating integrity as a shared value rather than a policing exercise. Clear, task-specific AI policies stated up front prevent most problems. Where genuine misconduct occurs, follow fair, transparent processes. A culture where students understand why their own thinking matters is far more effective than any detection tool.
Feedback and marking workload
Authentic and process-based assessment can be richer but also more time-consuming to mark, and workload is a legitimate concern that must be designed for rather than wished away. The goal is high-value feedback that improves learning, delivered sustainably. Front-load feedback where it changes outcomes — on drafts and process — rather than pouring detailed comments onto final products students will never revise. Use rubrics to make marking faster, more consistent and more transparent.
Batch common feedback, use audio or short video comments where they are faster and warmer than writing, and build peer and self-assessment into the process so students receive more feedback without all of it coming from you. Design your assessment schedule to spread marking across the term rather than concentrating it in unmanageable peaks. AcademicStaff's workload calculator helps you model the true marking load of a redesigned assessment before you commit to it, so authentic assessment enhances learning without breaking you. For sustaining yourself through busy marking periods, see our workload and wellbeing resources.
The AI-Resilient Assessment Redesign Toolkit
A free toolkit with an audit checklist, ten redesign patterns and a sample AI-use policy to help you convert vulnerable assessments into authentic tasks that measure genuine learning — without inflating your marking load.
Want the full article?
Enter your email for free access to the rest of this guide and our TEQSA resource library.
Frequently asked questions
Should I just go back to invigilated exams to beat AI?
Invigilated components have a role, but relying on them alone narrows assessment to what can be done under exam conditions and often tests recall over genuine capability. A better strategy is a diverse diet of authentic tasks that measure the human judgement, context and reasoning AI cannot supply well.
Is AI-detection software a reliable way to catch cheating?
No. Detection tools are unreliable, produce false positives that can wrongly accuse students — disproportionately those writing in a second language — and are quickly outpaced. Shift your strategy from detection to design: build assessments that make dishonesty difficult and pointless, and cultivate genuine integrity.
How do I stop authentic assessment from exploding my marking workload?
Front-load feedback onto drafts and process where it changes outcomes, use clear rubrics, batch common comments, try audio feedback, and build in peer and self-assessment. Spread marking across the term rather than in peaks, and model the true load before committing to a redesign.
The AcademicStaff editorial team writes practical, evidence-based guidance for university staff — drawing on sector reporting, funder guidelines and the lived administrative reality of academic work. Every guide is reviewed for accuracy against current Australian higher-education practice.
