Over the past year I’ve been working with a number of UK universities to rethink assessment for an AI-enabled world. While much of the public conversation has focused on AI policies and guidance, my work has centred on the assessment frameworks that underpin them: the decisions about what students are asked to do, how learning is evidenced, how work is marked, and how AI fits into that picture. Client confidentiality means I can’t share the artefacts that work has produced. But the thinking behind it, the critique of where the sector’s current approach goes wrong, and the practical alternative I now use with the institutions I work with, I believe is worth setting out on its own terms.
The trouble with AI scales
The dominant response to generative AI in assessment, almost everywhere I look, is some version of a scale: a numbered ladder, often colour-coded, telling students how much AI they are permitted to use on a given task. The best known is the AI Assessment Scale (AIAS), developed by Mike Perkins, Jasper Roe and Leon Furze, which has been adopted by hundreds of institutions worldwide since 2023 and is genuinely useful as a starting vocabulary. Its authors deserve credit for filling a real gap at speed.
What’s telling is that the AIAS’s own creators have since published a commentary on how their scale gets misused in practice. Writing in the Journal of Applied Learning & Teaching in 2025, Perkins, Roe and Furze catalogue the pitfalls they’ve observed across schools and universities: unenforceable “no AI” labels, scales bolted onto assessments without the assessments themselves being redesigned, equity blind spots, and what they call “control-first policies that encourage performance theatre.” Their central point is that the AIAS is a communication device, not an assessment security mechanism. Assigning an assessment an AIAS level does not make it secure, nor does it make it authentic in a world where AI is ubiquitous. If the assessment itself still rewards outputs that an AI system can readily produce, the stated level becomes little more than an expectation rather than a meaningful feature of the assessment. Authenticity has to be designed into the task itself, not declared through a label.
That maps closely onto the single biggest structural finding from the work I’ve been doing this year. Early iterations of the framework I was helping to develop categorised assessments according to the extent of AI use they permitted. On the surface this seemed neutral, but in practice it wasn’t. Staff and students naturally interpreted the categories as a hierarchy, with the middle position seen as a compromise rather than a valid assessment design in its own right. That ambiguity became the point at which inconsistent marking, unclear disclosure expectations and inequitable outcomes were most likely to emerge. Reframing the framework around distinct assessment approaches, or assessment ‘ecosystems’ as I personally prefer, each with its own design principles, removed the implied hierarchy and made design, rather than restriction, the organising principle.
More fundamentally, the problem is that any scale, however well designed, answers the wrong first question. It tells you what a student may or may not do. It does not tell you what the task is meant to develop, or how you would know whether a student can do it. Permission is downstream of purpose, yet most frameworks begin with permission.
What a university is actually for
It is worth returning to a plainer point, because it’s easy to lose it in the operational detail of policy and process: the purpose of a university education is to teach students how to think — to reason from evidence, interpret ambiguous material, evaluate competing arguments, solve problems that don’t have a single correct method, and generate ideas that are genuinely their own. Generative AI changes how information and text get produced. It does not change that underlying purpose, and it shouldn’t be allowed to quietly displace it.
This matters for how assessment gets designed, not just how it gets policed. If you start from ‘what can a student do with AI on this task,’ you inherit a compliance frame, and the whole apparatus you build, declarations, detection, disciplinary process, follows from that. If you start from ‘what thinking does this task need to build and evidence, and does AI help or substitute for that,’ AI becomes one design variable among several rather than the organising question. It’s a small reframing with large downstream consequences: it changes what gets assessed (process and judgement, not only the final output), what evidence gets collected (a trail across a programme, not one high-stakes artefact), and what gets named explicitly as a marking criterion, critical, independent use of AI and judgement about when not to use it, assessed as a capability in its own right rather than treated as an afterthought to a permission statement.
PACIER: giving ‘critical thinking’ some structure
Higher education talks about critical thinking constantly and defines it inconsistently. The word is frequently used in course learning outcomes, but rarely is the term made explicit enough to design an assessment against. One framework I’ve found useful for giving that conversation some shared vocabulary is PACIER, set out by Hyo Jeong Shin and colleagues in the Journal of Intelligence in 2025 (building on work associated with Macat International and collaborating institutions). PACIER describes critical thinking as six interrelated facets: Problem-solving, Analysis, Creative thinking, Interpretation, Evaluation and Reasoning.
I don’t treat PACIER as a checklist to bolt onto a rubric. Its value is as a design lens: instead of asking whether a student’s AI use was ‘appropriate,’ you ask which of those six facets a given task is actually meant to develop, what a well designed task for that facet should include, and whether the way AI is available to students in that task supports the student doing that thinking themselves, or lets them route around it. That is a sharper, more answerable design question than any permission scale gives you, and it travels well across disciplines that otherwise struggle to agree on a shared definition of rigour.
A framework outline
I can’t share the specific descriptors, rubrics or exemplars built for any of the institutions I have worked with — that work belongs to them. But the underlying method is transferable, and it’s the one I now bring to every assessment redesign conversation:
Audit before you design. Map what current assessment already asks of students against something like PACIER’s facets. Most programmes are heavy on recall and analysis and light on evaluation and creative thinking — which is exactly where AI creates the most displacement risk.
Categorise by structural intent, not permission level. Ask what kind of assessment condition a task represents. Is it a controlled, unassisted demonstration where AI use can be prohibited by the assessment environment; or a task where AI is a legitimate collaborator, or a task where AI use is itself the object of study. Naming these as categories or ecosystems, not levels, removes the implied hierarchy and forces each one to be independently well designed.
Redesign the brief and the rubric together. A permission statement changes nothing on its own; the evidence you ask for and the criteria you mark against have to change with it, including making critical, independent use of AI an explicit, assessable criterion rather than a footnote. And don’t underestimate the importance of being explicit and clear to students what can they typically use AI for, and what areas will impact their grade if they use it as a substitute for the thinking and skills they need to evidence?
Build a trail, not a moment. Evidence of a student’s thinking accumulates more reliably across drafts, process logs and staged submissions than it does in a single high-stakes artefact, a principle that both PACIER-style thinking models and the AIAS’s own recent guidance converge on independently. AI Declaration statements are useful as one point in a richer evidential trail that allows markers to see how a student’s thinking developed over time.
Train the markers, not just the students. Most of the inequity in this space isn’t about how students use AI. It’s about inconsistent staff interpretation of ambiguous guidance, applied at scale across a large programme or institution.
Beyond the scale
The scales that spread through the sector in 2023 and 2024 did something genuinely useful: they got AI onto assessment agendas quickly, at a moment when most institutions had nothing. But a scale is a communication device, not an assessment framework, and the sector is far enough into this now that the distinction matters. The institutions I see getting real value from this work are the ones treating AI as one more input into a much older question: how does this task actually ask a student to think, and how do we know they can. Get that question right, and the AI question becomes considerably smaller than it currently feels.
Putting it into practice
I try to practise what I preach. Alongside my strategic HE consulting, I’m also a Subject Lead for a portfolio of MA Education programmes, where I’ve been working with our team to rethink assessment for an AI-enabled world. That means designing assessments that are engaging, fair and authentic; helping students understand not just what is expected of them, but why an assessment has been designed in a particular way and how they can get the most from it; and developing rubrics that recognise the qualities AI cannot demonstrate, but our students can.
If I’m honest, this is the part of my work that makes me lose track of time. I’m fascinated by assessment design because the impact is almost immediate. You can see students engaging differently, thinking more deeply and producing work that genuinely reflects their own judgement. I love reading and grading submissions that carry the originality and insight that the assessment seeks because it means I learn from the students too. As much as I enjoy the strategic work I do with universities (and I really do), education has always been at the heart of what motivates me. Ultimately, my job isn’t to teach students content or what to think, but how to think and how to expess that.
If you’re working through what this looks like in your own institution or programme, I’d be delighted to talk it through.