← All insights

Assessment in the Age of AI

A practical guide to designing assessment that embraces AI while making student understanding, judgement and adaptability visible.

Assessment evidenceLive model
THE QUESTION HAS CHANGEDWhat can this person understand, judge and do?
01 / PRODUCTSubmitted work

Useful evidence, but no longer sufficient alone.

02 / PROCESSDecisions made

Sources, drafts, checks and declared AI use.

03 / PERFORMANCECapability seen

Explanation, judgement and adaptation.

Confidence increases through convergenceHuman capability visible
The old proxyA polished product implied a capable person.

That inference once carried much of the assessment judgement.

The new ambiguityExcellent output can hide uneven understanding.

AI can improve the artefact without improving every human capability behind it.

The design responseGather evidence from more than one moment.

Combine making with explanation, independent judgement and critique.

A practical guide to making human capability visible

A student submits a thoughtful report. It is well structured, carefully written and supported by credible evidence. The work may be excellent. Yet one question now sits behind every polished submission:

What does this work tell us about the capability of the person who submitted it?

That question is not an accusation. It is a design problem.

Generative AI can help someone plan, research, analyse, write, code, illustrate and revise. These are valuable capabilities, and students should learn to use them well. But the finished product can no longer carry the whole burden of proof. A strong submission might show that a student can direct AI effectively. It might also hide gaps in their understanding. Usually, the product alone cannot tell us which is true.

The answer is not to pretend AI can be removed from education. Nor is it to turn every assessment into an exercise in surveillance. The more useful response is to create a small number of purposeful moments in which human capability becomes visible.

AI can improve the product without improving the person to the same degree.

Use the model that follows to redesign a single task, a module or a whole programme. It gives lecturers, course leaders and quality teams a common language for deciding what a qualification should genuinely certify.

The shift from product to evidence

For a long time, the submitted product acted as a reasonable proxy for the work behind it. If a student produced a strong essay, report, presentation or piece of code, we inferred that they understood the subject and could perform the thinking involved.

That inference is now less secure.

AI can improve the product without improving the person to the same degree. It can smooth weak prose, propose a convincing structure, generate plausible analysis and supply an answer that sounds more certain than the underlying evidence allows. It can also help a capable student work faster, explore alternatives and produce better work. The same visible output can therefore represent very different levels of human understanding.

Assessment has to gather a richer set of evidence.

The evidence stack Confidence comes from convergence, not one perfect artefact
01 / PRODUCT

What was made

The finished work shows quality, relevance and the ability to deliver.

  • Report or artefact
  • Research or design
  • Code or performance
  • Final recommendation
02 / PROCESS

How it developed

Selected traces reveal choices, checks and responsible use of tools.

  • Plans and drafts
  • Source decisions
  • Material AI use
  • Revision after feedback
03 / PERFORMANCE

What the person can do

Observed performance makes understanding and judgement directly visible.

  • Explain a choice
  • Judge new evidence
  • Respond to challenge
  • Critique an AI answer
One signalTriangulated confidence

The question is no longer simply, “Did this student produce this document?” It is:

What evidence would allow us to certify that this student possesses the capability represented by this qualification?

Once that question is asked clearly, assessment can be designed backwards from it.

Begin with the capability, not the task

Before rewriting an assignment, name what the student must be able to do. Broad phrases such as “demonstrate understanding” are too vague to guide design. An assessment team needs to decide whether understanding means recalling essential knowledge, interpreting evidence, applying a method, making a professional judgement or explaining a decision to another person.

The distinction matters because these capabilities are not interchangeable. A student may remember the relevant concepts yet struggle to use them in a case. Another may reach a sound conclusion but be unable to explain the evidence behind it. A third may perform well when the situation is familiar but miss the moment when a new risk makes the original answer unsafe.

AI adds another capability to the set: the ability to use a powerful tool while remaining accountable for the result. This is not demonstrated merely by producing a better document. It becomes visible when the student can recognise weak evidence, reject a plausible error and explain why the final decision remains theirs.

A student can therefore use AI to improve a report and still be required to explain its argument. They can use a coding assistant during a project and still be asked to diagnose an unfamiliar fault without it. They can use AI to explore a case and still be expected to make a safe decision when new facts arrive. The assessment method should change with the capability being claimed.

The goal is not to prove that every word came from the student. The goal is to observe enough of the student’s thinking to make a credible judgement about their capability.

The six stage model

The model has six connected stages. Each adds a different form of evidence.

The six stage model One assessment journey, three conditions for AI
AI enabledAI restrictedAI returns
01CREATEMake

Produce meaningful work and show how it developed.

02OWNExplain

Show understanding of important choices.

03DECIDEJudge

Reach a reasoned conclusion independently.

04RESPONDChallenge

Handle information that was not rehearsed.

05REVISEAdapt

Change position when the evidence requires it.

06CONTROLCritique

Use AI without surrendering judgement.

THE COMPLETE CLAIM

The student can make, understand, judge, respond, revise and remain accountable when AI is present.

Each stage answers a different question. Together, they give a fuller picture than a final submission can provide on its own.

The six stages do not need to become six separate assessments. In many cases, a project, a short offline exercise and a brief conversation are enough to cover the whole model.

Stage one: Make

Can the student develop meaningful work?

The student creates something substantial: a report, design, investigation, performance, campaign, programme, product, case analysis or research project. This remains an important part of assessment. A finished piece of work shows whether the student can bring knowledge, method and communication together in a form that another person can use.

AI may be permitted because using it can be part of authentic contemporary practice. The institution should state what assistance is allowed, what must be declared and what remains the student’s responsibility. A vague instruction to “acknowledge AI” is not enough. Students need to know whether the relevant unit of disclosure is a tool, a prompt, a passage or a material decision.

The final product should be accompanied by a small amount of selected process evidence. Early plans can show how the problem was framed. A draft or version history can show meaningful development. A note about sources, methods and feedback can show where judgement was exercised. If AI materially influenced the direction, analysis or expression of the work, that influence should be described.

The word selected is important. An assessor does not need every click, prompt or minor revision. Excessive process capture creates noise and encourages performance for the record. Ask for evidence that answers four substantive questions: where AI influenced the work, which suggestions were accepted or rejected, how important claims were checked and which consequential decisions remained with the student.

Framed in this way, disclosure is not a confession. It is an account of professional responsibility.

Evidence question: What did you produce, and how did you get there?

Stage two: Explain

Does the student understand their own work?

A short conversation can reveal more than another thousand words of written reflection. The purpose is not to make the student repeat the submission. It is to discover whether they understand the choices on which it depends.

The most revealing questions are specific. An assessor can point to a paragraph, chart, design choice, source or block of code and ask why that approach was chosen. A second question can explore the alternative that was rejected. A third can ask which evidence carries the greatest weight or which part of the argument is weakest. Asking the student to explain a technical section in simpler language often reveals whether they possess the idea or only its vocabulary.

The conversation should feel intellectually serious without becoming adversarial. A calm exchange gives the student room to think, qualify and correct themselves. The standard is not instant fluency. It is ownership of the reasoning.

This stage can be a viva, a tutorial conversation, an audio explanation or a recorded walkthrough. Where a live conversation would create an unnecessary barrier, an accessible alternative can preserve the same evidential purpose.

Evidence question: Can you explain what you did and why?

Stage three: Judge

Can the student reach a reasoned conclusion independently?

At this stage, introduce a short, supervised assessment window in which generative AI and internet access are absent. The event should be open book wherever recall is not the capability being tested.

Students might receive their own notes, selected readings, a dataset, a case file, a formula sheet or reference documents. They then use those materials to analyse a problem and make a judgement.

The task should not reward memory by accident. If the intended capability is evaluation, provide the relevant information and ask the student to decide what it means. A useful judgement task requires a choice, competing considerations and enough uncertainty for reasoning to matter.

For example, a student might recommend one of three options and identify the evidence that is decisive. They might interpret findings that support more than one conclusion, rank risks that cannot all be addressed at once or decide what action is justified while acknowledging what remains unknown. The final question should make responsibility explicit: what would you do next if you were accountable for the outcome?

The event can be handwritten, completed on a locked device or conducted as supervised practical work. The format matters less than the evidence it produces.

Evidence question: Can you reach and defend a judgement when AI is not doing the thinking for you?

Stage four: Challenge

Can the student handle something they did not anticipate?

After the student has formed a position, change the situation. The new information should matter to the judgement already made. A source might contradict the central claim. A key assumption may prove false. A budget may be reduced, a stakeholder may object, a dataset may change or a safety concern may emerge. In practical subjects, the preferred component, treatment, argument or course of action might simply fail.

Then ask one simple question:

What changes now, and why?

This is one of the strongest moments in the model because the response could not have been fully prepared in advance. It does more than authenticate the original work. It shows whether the student can reason when the script breaks.

The challenge should be meaningful but proportionate. Its purpose is not to catch the student out. It is to reveal how they respond when reality becomes less tidy than the original task.

Evidence question: Can you reason under uncertainty?

Stage five: Adapt

Can the student revise their judgement intelligently?

The student now revisits the original position. A strong response identifies what must change and what should be retained. It names the evidence that caused the revision, distinguishes facts from assumptions and states what further information would be sought before acting.

Changing an answer should not automatically be treated as weakness. In professional life, refusing to change in the face of better evidence is often the greater failure.

Assessors should distinguish between an unsupported reversal and a thoughtful revision. The first abandons a position. The second shows that the student understands which parts of the reasoning still hold and which no longer do.

This stage makes intellectual flexibility visible. It also allows students to demonstrate mature judgement even when their first answer was imperfect.

Evidence question: Can you change your position for good reasons?

Stage six: Critique

Can the student judge AI rather than simply use it?

Now bring AI back deliberately.

Give the student an AI generated response, or allow them to generate one. Ask them to examine it as a fallible contribution rather than an authority. They should identify unsupported claims, expose hidden assumptions and verify facts that would materially affect the decision. They should also notice what is absent: an overlooked perspective, a missing source, a practical constraint or an uncertainty disguised by fluent language.

The student then compares the response with their own reasoning. They decide what can be retained, what must be rejected and how the answer should be improved. The final question is about accountability. If this recommendation causes harm or fails, who was responsible for accepting it?

This stage treats AI literacy as a form of judgement, not a collection of prompting tricks. Producing fluent text with AI demonstrates tool use. Recognising when that text is misleading demonstrates critical capability.

Evidence question: Can you attain the decision maker when AI is available?

Where AI belongs Match the tool condition to the capability being assessed
Assessment purposeAI conditionUseful evidence
Explore and produce

AI use reflects authentic contemporary practice.

Permit
Product plus concise process record

Important prompts, checks, source choices and revisions.

Show understanding

The assessor needs to hear the student’s own account.

Restrict
Explanation anchored in the submission

Why this choice, what alternative, where is the weakness?

Reach an independent judgement

The capability must be demonstrated without generated reasoning.

Remove
Open book offline decision

Interpret supplied evidence and defend a conclusion.

Respond to change

The answer cannot be fully prepared in advance.

Remove
Unseen challenge and revision

Show what changes, what remains and why.

Evaluate AI

Responsible tool use is itself part of the capability.

Return
Critique, verification and improvement

Identify error, missing evidence and unsafe confidence.

Creating a useful AI restricted window

Not every assessment needs to be completed without AI. A programme can create short, deliberate verification windows at the point where independent human capability matters most.

The format should follow the capability. Written analysis can be observed through a short open book paper case. Conceptual understanding may be better seen in a ten minute viva or whiteboard explanation. Technical reasoning may require a supervised coding or problem solving task. Professional judgement may be best demonstrated through an unseen case followed by a decision and defence.

Handwriting is not inherently more authentic, and live speech is not inherently more rigorous. Each format is useful only when it gives the student a fair opportunity to demonstrate the intended capability.

Four design choices make these windows more useful.

Make the purpose explicit

Tell students which capability needs to be demonstrated independently and why. “No AI” is a rule. “This activity allows you to demonstrate your own clinical reasoning” is an assessment rationale.

Keep the window proportionate

The window only needs to produce enough evidence for a confident judgement. A focused fifteen minute conversation may be more useful than a second full examination.

Provide the knowledge students need

If recall is not being assessed, permit notes and provide source material. Removing the internet does not require removing useful reference documents.

Mark the thinking, not the polish

Work produced under time pressure will look different from a refined submission. The rubric should reward interpretation, reasoning, priorities, justified decisions and recognition of uncertainty. Surface polish should carry little or no weight unless communication under pressure is itself a required capability.

The objective is not surveillance. It is to create sufficient independent evidence for the institution to say:

We have observed this student’s capability, not merely their output.

One architecture, not six extra assessments

The model becomes burdensome if every stage is turned into a separate submission. A better approach is to combine stages within one coherent assessment.

Choose the light model when the existing task is sound and only needs stronger verification. Choose the integrated model when professional reasoning is central. Choose the programme model when the award needs broad evidence without repeating the same event in every module.

Three workable architectures Choose the smallest pattern that produces credible evidence
MODEL 01 / LIGHT

Project plus verification

Keep the existing project. Add a short conversation and one unseen question.

MAKEEXPLAINTEST
Best for
Existing modules
Live time
10 to 15 minutes
Evidence gain
Ownership and judgement
MODEL 02 / INTEGRATED

Project plus live session

Combine explanation, an unseen change and AI critique in one structured event.

MAKEVERIFYADAPT
Best for
Professional judgement
Live time
20 to 30 minutes
Evidence gain
Depth and adaptability
MODEL 03 / PROGRAMME

Evidence across modules

Distribute independent analysis, practical performance and AI critique across the award.

BUILDOBSERVECRITIQUE
Best for
Large programmes
Live time
Shared across modules
Evidence gain
Qualification level claim

A worked example: business strategy

Assume the module outcome requires students to evaluate strategic options, make an evidence based recommendation and defend that recommendation when conditions change. A conventional assessment might ask for a strategy report. The report would show whether the student can assemble a persuasive case, but it would reveal less about ownership, transfer and judgement.

The worked example below uses a fictional commercial furniture manufacturer called Alder Works. The board can invest no more than £2 million and must choose among three routes to growth. The example follows a student who initially recommends a furniture refurbishment service.

The six stages are not six separate assignments. The report is completed first. Explanation, independent judgement, challenge, adaptation and critique are then combined within one structured verification event. The same strategic decision runs through the whole assessment, so each stage adds evidence rather than starting the student again.

Worked assessment One strategic decision examined through six forms of evidence
ILLUSTRATIVE CASE / ALDER WORKS

How should the business grow without weakening the business it already has?

Alder Works is a fictional commercial furniture manufacturer. The student acts as an adviser to its board and must recommend one credible route for the next three years.

REVENUE£24 millionCurrent annual revenue
OPERATING MARGIN8 per centLimited room for error
CONCENTRATION28 per centRevenue from one customer
INVESTMENT LIMIT£2 millionAvailable over three years
THE BOARD DECISION

Recommend one route: pursue public sector contracts, launch a furniture refurbishment service or enter a new European market through a distributor. Compare the options, reject two and state the conditions under which the chosen route remains viable.

01Make
Prepare the board paper

The student recommends the refurbishment service. AI is permitted for research, scenario exploration and editing. Material use is declared.

A defensible strategic choice

The recommendation links customer demand, operational capacity, cash exposure and concentration risk. Assumptions are visible and the rejected options receive fair consideration.

02Explain
Defend two consequential choices

The assessor selects the demand estimate and the decision to build internal refurbishment capacity. The student explains the evidence, alternatives and uncertainty.

Ownership of the reasoning

The student can distinguish evidence from inference, explain why the export option was rejected and identify the demand assumption most likely to fail.

03Judge
Analyse an unseen evidence pack

Without AI, the student receives revised unit costs, capacity data and customer research. They decide whether the service still meets the board’s investment test.

An independent decision rule

The answer identifies the minimum demand needed, the cash exposure the business can tolerate and the evidence that carries most weight. Uncertainty is not disguised.

04Challenge
Introduce a material shock

The largest customer will leave in six months. The business now faces a potential revenue loss of £6.7 million and cannot safely commit the full investment.

Priorities change immediately

A credible response recognises that cash protection and customer concentration now take precedence over rapid expansion. Repeating the original plan is no longer defensible.

05Adapt
Revise the route, not just the wording

The student retains the refurbishment proposition but delays the facility, tests demand through a delivery partner and creates a lower cost pilot.

A coherent revision

The revised plan protects working capital, sets a cash floor and requires three anchor customers before further investment. The original market evidence still supports a controlled experiment.

06Critique
Evaluate an AI recommendation

The supplied AI response proposes a twenty per cent discount and immediate construction of the facility to replace lost revenue quickly.

Judgement remains with the student

The student rejects the unsupported demand assumption, tests the effect on margin and identifies the cash risk. They retain only the useful suggestion to accelerate customer interviews.

ORIGINAL POSITIONInvest £1.6 million in an internal refurbishment service.
REVISED POSITIONRun a partner delivered pilot before committing fixed capital.
WHY THE CHANGE IS STRONGThe strategy responds to new cash risk while preserving the part of the market thesis still supported by evidence.

A weak response to the customer loss would simply reduce the budget or repeat the original recommendation more cautiously. A strong response recognises that the purpose and sequence of the strategy must change. It protects working capital, delays fixed investment and converts the proposal into a controlled pilot with explicit evidence gates.

The strongest response also knows what should not change. The customer research may still support demand for refurbishment. Abandoning the idea entirely would ignore valid evidence. Good adaptation preserves the part of the original thesis that remains sound while changing the commitment that the business can no longer afford.

The AI critique completes the example. The student is not rewarded for disagreeing with AI as a matter of principle. They are rewarded for identifying the unsupported demand assumption, calculating the margin consequence and retaining the one suggestion that is useful. This is what it means for the human to remain the decision maker.

How the model travels across disciplines

The same structure can be adapted without making every subject look the same.

Across disciplines Keep the model, change the authentic form of evidence
COMPUTING

Diagnose and defend

Build with AI, explain an architectural choice, diagnose an unseen fault and review an AI patch.

Visible capability: reasoning about code, not merely producing it.
HEALTH

Prioritise safely

Prepare a case, explain priorities, respond to a changed condition and find unsafe AI assumptions.

Visible capability: safe judgement when information changes.
LAW

Interpret authority

Draft advice, defend source selection, respond to a new precedent and critique an overstated AI summary.

Visible capability: justified interpretation and professional caution.
ENGINEERING

Work within constraints

Develop a design, explain tradeoffs, respond to component failure and assess an AI redesign.

Visible capability: technical judgement under real constraints.
CREATIVE PRACTICE

Direct with intent

Develop work with appropriate tools, explain creative choices and adapt to an unseen brief.

Visible capability: purposeful direction rather than generic output.
HUMANITIES

Build an interpretation

Use evidence, analyse an unseen source and expose omissions in an AI generated argument.

Visible capability: interpretation that survives challenge.

Discipline matters. The framework stays consistent, but credible evidence must reflect the real forms of judgement used in the field.

Design the marking criteria around visible capability

A conventional rubric often gives most of its attention to the final product. That made sense when the product could carry most of the evidential burden. In the present context, the rubric should recognise a wider range of observable qualities without fragmenting the assessment into dozens of tiny criteria.

One part of the judgement should still concern the quality and relevance of the finished work. A second should concern understanding: whether the student can explain important choices and evaluate the evidence used. A third should address independent judgement, including the response to new information and the ability to revise a position coherently. Where AI use is permitted, the rubric should also recognise critical use, appropriate verification and honest explanation of limitations.

Not every criterion needs a separate mark. Some may be threshold requirements. For example, a professional course might require students to demonstrate safe independent judgement even if the final project is excellent.

Illustrative marking profile Move some weight from polish to observable capability

A balanced evidence profile

This is an example, not a universal formula. Adjust the balance to the discipline and any threshold capabilities required by the award.

Finished work35 per cent
Understanding and process20 per cent
Independent judgement20 per cent
Challenge and adaptation15 per cent
Critical use of AI10 per cent
Professional programmes may also require safe independent judgement as a threshold, regardless of the total mark.

The rubric should also make clear that confidence is not the same as competence. A hesitant student may reason carefully. A fluent student may be confidently wrong. Assess the quality of the thinking that can be observed.

Accessibility is part of validity

An assessment is not more human simply because it happens live or on paper. Any verification method can introduce barriers that have little to do with the capability being tested.

A viva may disadvantage a student with a communication related disability. Handwriting may create a barrier for someone who normally uses assistive technology. A tightly timed event may measure processing speed when the course actually intends to measure judgement.

The solution is to preserve the evidential purpose while adapting the format. If the purpose is to observe reasoning, the student may be able to demonstrate it through an accessible locked device rather than handwriting. Questions can be provided in writing as well as spoken aloud. Additional time, rest breaks or a quieter room can remove barriers without changing the intellectual demand.

Sometimes the mode itself should change. A recorded explanation may provide equivalent evidence to a live response. In another case, a structured written commentary may demonstrate the same ownership as a viva. The test is not whether every student completed an identical activity. The test is whether each student produced comparable evidence of the intended capability.

Reasonable adjustment does not weaken verification. It removes irrelevant obstacles so the intended capability can be seen more accurately.

Manage workload by sampling intelligently

The most common practical concern is staff time. Short conversations and supervised events can appear modest until they are multiplied across a large cohort. A twelve minute conversation with one hundred and twenty students requires twenty four hours of assessor time before transitions, administration and moderation are included. The design must therefore be economical as well as valid.

Economy comes from sampling. The assessor does not need to discuss the whole submission. Two or three consequential decisions are usually enough to test ownership and understanding. Explanation, challenge and adaptation can occur in the same conversation: the student explains a choice, receives one changed condition and revises the position.

Consistency does not require identical questions. A shared bank can define the type and difficulty of question while each prompt remains anchored in the student’s work. One assessor might ask about a rejected source, another about an assumption in a model, but both are testing the ability to justify a consequential choice.

For larger cohorts, programmes can use structured group sessions in which the stimulus is shared but each student records an individual judgement. They can also place substantial verification at programme level rather than repeat it in every module. A concise assessor record and moderation of a representative sample are normally more useful than a full transcript of every exchange.

The aim is sufficient evidence, not exhaustive observation.

What to tell students

Students need a clear account of where AI is permitted, where it is restricted and why. Ambiguity creates anxiety and rewards guesswork.

The brief should name the AI uses allowed in the main task and explain what material use must be declared. It should identify any process evidence the student needs to retain, rather than announcing that requirement after submission. It should describe the part completed without AI, the capability that part is designed to assess and the resources that will be available.

Students should also see how the different forms of evidence contribute to the result. If the viva can change a mark, confirm a threshold or trigger further review, that consequence must be stated. Accessibility arrangements should be presented as part of the assessment design, not as an exception hidden elsewhere.

The tone matters. Students should understand that the process is designed to let them demonstrate what they know, not to lure them into an integrity violation.

One useful formulation is:

You may use AI during the project within the stated rules. You remain responsible for every claim and decision in your submission. The follow up activity gives you an opportunity to demonstrate your own understanding, judgement and ability to respond to new information.

Common design mistakes

Treating authentication as the whole problem

Attempts to identify AI use from writing style are weak substitutes for assessment evidence. Fluent prose may be assisted, heavily edited or entirely unaided. The surface alone rarely establishes the process with confidence. Blanket bans create a similar problem when they are not tied to a capability. They state what students must not do without clarifying what the institution needs to observe.

Excessive process capture does not solve this weakness. Thousands of prompts and screenshots produce a large record but little insight. The assessor needs concise evidence about consequential decisions, verification and revision. The purpose is to understand the student’s relationship to the work, not to reconstruct every action taken while producing it.

Confusing pressure with rigour

An adversarial viva can measure confidence under pressure more strongly than understanding. A closed book task can measure memory even when the intended outcome is analysis. Neither is necessarily rigorous merely because it is difficult.

Rigour comes from alignment. If students need information to make a judgement, provide it. If they need time to interpret a complex source, allow it. Use clear prompts and predictable criteria, then place the difficulty in the quality of reasoning required.

Rewarding defence instead of judgement

Students should not feel compelled to preserve an original answer after the evidence has changed. That rewards stubbornness and can teach the wrong lesson about expertise. A well designed assessment gives credit for identifying precisely when a position no longer holds.

The same principle applies to the assessment itself. Verification should replace or streamline weaker activity wherever possible. Adding a viva, a challenge and an AI critique without removing anything creates volume rather than validity.

A practical redesign process

Course teams can use the canvas below to move from a qualification claim to a workable assessment architecture.

Assessment design canvas A one page conversation for course teams
Make human capability visibleCOMPLETE FROM LEFT TO RIGHT
01 / CLAIM

What can a successful student do?

Write one observable capability, not the name of the assignment.

02 / AI ROLE

Where is AI authentic?

Name where AI use is valuable and where it would hide essential performance.

03 / EVIDENCE

What will make capability visible?

Select product, process and observed performance evidence.

04 / CHANGE

What new fact will test adaptation?

Use one realistic change that matters to the original judgement.

05 / CRITIQUE

How will AI return?

Ask the student to find error, missing evidence and misplaced confidence.

06 / BURDEN

What can be removed?

Replace weak evidence. Do not add verification on top of everything else.

Define the claim and locate AI

Complete the sentence, “A student who passes this assessment can…” Use an observable capability rather than the name of the task. Then examine the current assignment honestly. Identify what AI can already complete well and what still requires contextual, practical, ethical or interpersonal judgement.

Mark the difference between evidence that is visible and capability that is merely inferred. A strong report makes quality visible. It does not automatically make independent judgement visible. This gap tells the team where verification is needed.

Design the evidence sequence

Choose the smallest verification point that strengthens a material claim. It may be a viva, unseen case, supervised practical, live explanation or offline judgement. Add one realistic change that requires the student to reconsider the original reasoning, then bring AI back through a critique connected to the same capability.

Read the sequence as one assessment experience. Each part should contribute evidence that another part cannot provide more efficiently. If two activities reveal the same thing, remove the weaker one.

Align judgement and communication

Revise the rubric so it rewards the capabilities now being observed. Decide whether any capability is a threshold rather than a source of additional marks. Write descriptors that distinguish a justified judgement from an assertion, and a coherent revision from an unsupported change of mind.

Explain the architecture before students begin. They should understand the role of AI, the purpose of the verification window and the basis on which their responses will be judged.

Test validity and burden

Estimate student time, assessor time, room needs, technology and accessibility requirements. Check whether the assessment can be delivered consistently at the actual cohort size. Remove activity that no longer contributes useful evidence.

Finally, ask colleagues to review the logic. Can they see a clear connection between the capability claimed, the task set, the evidence produced and the judgement made? If any link is weak, redesign that link before adding more assessment.

Questions for programme teams

Before approving a redesigned assessment, return to the claim made by the qualification. What human capability is being certified, and where can it be observed directly? If the answer still depends entirely on the quality of a submitted product, the evidence remains too narrow.

Next, examine the role of AI. Its use should be permitted where it is authentic and educationally valuable. It should be restricted where it would prevent the programme from seeing essential independent performance. Any restricted window should be long enough to produce credible evidence, but no longer.

Then review the intellectual demand. The challenge should test adaptability rather than surprise for its own sake. Students should receive credit for changing their mind when evidence warrants it. The critique stage should test judgement about AI, not merely familiarity with a tool.

Finally, examine the assessment as a system. The arrangements should be accessible, proportionate and possible to deliver consistently. An external reviewer should be able to follow the chain from the capability claimed, through the evidence observed, to the judgement made about the award.

The principle to keep

The central challenge is not proving that AI was absent from every part of a student’s education. That goal is neither realistic nor desirable.

The challenge is designing enough contrast to see the person clearly.

Let students use AI to explore, build and improve. Create deliberate moments in which they must explain, judge and adapt without it. Then bring AI back and ask them to question it with the same seriousness they would apply to any other source.

Do not try to remove AI from education. Create enough moments without AI to see what the human can do, and enough moments with AI to see whether the human remains in control.

Rapid Innovation Lab

Design for evidence, not suspicion.

Make human understanding, judgement and responsible AI use visible.

Start a conversation