When the Grade Becomes the Goal
When Assessment Stops Measuring Learning.
A university had begun noticing a growing pattern. Students were becoming increasingly dependent on AI when completing major assignments. The institution's first response was to look at AI use. But one faculty member asked a different question: What if the problem was not only what students were doing with AI, but what the assessment system was asking them to do in the first place?
Read the case →
CASE STUDY 04
Educational Management & Learning Design
The Case
A university had begun noticing a change in student behaviour.
Assignments were becoming increasingly polished. The language was sophisticated. The structure was often excellent. Students were producing work that appeared more professional than some of their previous submissions.
At the same time, faculty members were becoming increasingly concerned about generative AI.
Some students openly acknowledged using AI to improve their writing. Others used it to generate ideas, explain difficult concepts or restructure their work. Some appeared to be using it much more extensively.
The institution therefore began discussing familiar questions:
- Should students be allowed to use AI?
- How much AI assistance is acceptable?
- How should students declare AI use?
- How can teachers identify inappropriate use?
- Should assessments be redesigned?
All of these were reasonable questions.
But one faculty member asked a different one:
What exactly are our assessments measuring?
The question created some discomfort.
Because if the answer was simply “the quality of the final submission,” then the problem might be larger than AI.
It might be a problem with assessment itself.
--------------------------------------------------------------------------------------------------------
The ProblemThe university's immediate problem appeared to be AI-assisted assignments.
But the deeper problem was more difficult.
The university's immediate problem appeared to be AI-assisted assignments.
But the deeper problem was more difficult.
Traditional assessment often depends heavily on the final product.
A student submits an essay, a report, a presentation, a project, a piece of code, a design.
The teacher evaluates the result and assigns a grade.
But what happens between the beginning and the final submission?
- The student's uncertainty.
- The mistakes.
- The questions.
- The revisions.
- The feedback.
- The moments when understanding changes.
- The struggle to solve a problem.
- The decisions made along the way.
Much of that may disappear. What remains is the final product.
And that creates an important question:
If two students produce equally impressive final products, do we necessarily know that they have demonstrated the same learning?
The answer depends on what the assessment was designed to measure.
This is where the university began looking more closely at the relationship between learning outcomes, learning activities and assessment.
Constructive alignment argues that these elements should work together: what students are expected to learn should be reflected in the activities they undertake and in how their learning is assessed.
But the university realised that alignment was not simply a curriculum-design issue.
It was becoming an AI issue too.
A student submits an essay, a report, a presentation, a project, a piece of code, a design.
The teacher evaluates the result and assigns a grade.
But what happens between the beginning and the final submission?
- The student's uncertainty.
- The mistakes.
- The questions.
- The revisions.
- The feedback.
- The moments when understanding changes.
- The struggle to solve a problem.
- The decisions made along the way.
Much of that may disappear. What remains is the final product.
And that creates an important question:
If two students produce equally impressive final products, do we necessarily know that they have demonstrated the same learning?
The answer depends on what the assessment was designed to measure.
This is where the university began looking more closely at the relationship between learning outcomes, learning activities and assessment.
Constructive alignment argues that these elements should work together: what students are expected to learn should be reflected in the activities they undertake and in how their learning is assessed.
But the university realised that alignment was not simply a curriculum-design issue.
It was becoming an AI issue too.
-----------------------------------------------------------------------------------
When the Grade Becomes the Goal
Students are not indifferent to grades. They can't afford to be. In this highly competitive socio-economic environment grades can determine progression, scholarships, opportunities, future study and even long-term success.
So when a major assignment carries significant weight, it is understandable that students pay close attention to the outcome.
But something subtle can happen.
The student's question can gradually change from:
What can I learn from this assignment?
to:
What do I need to do to get a good grade?
The two questions are not the same.
And once the grade becomes the primary target, students may naturally begin looking for the most efficient way to produce the required result.
That does not necessarily mean they are dishonest. It may simply mean they are responding rationally to the environment in which they are being assessed.
- If the system rewards the final product, students will pay attention to the final product.
- If the system rewards memorisation, students will devote time to memorisation.
- If the system rewards polished written work, students will try to produce polished written work.
The assessment system does not merely measure behaviour. It can also shape behaviour.
That observation has long been part of the wider discussion about assessment and its consequences.
Assessment validity is not only about whether a task produces evidence of achievement; it also involves considering the consequences and effects of assessment practices.
The university therefore began asking:
What behaviour is our assessment system encouraging?
-----------------------------------------------------------------------------------
The Polished Answer
Consider two students.
Student A
Student A understands the subject but struggles with academic writing. His first draft contains awkward sentences. His explanation is not particularly elegant. He makes several mistakes.
But the work shows their own reasoning. He has clearly wrestled with the problem.
Student B
Student B understands some parts of the subject. He is less confident about the underlying concepts. But he uses AI to improve the structure and language of his assignment. The final submission is fluent, well organised, confident, sophisticated.
If the assessment primarily rewards the final product, Student B may appear to have demonstrated stronger learning.
But what does the final submission actually tell us?
It tells us something ... but perhaps not everything we think it tells us.
We may know that the final answer is good.
- We may not know how much of the reasoning belongs to the student.
- We may not know what the student initially understood.
- We may not know what they struggled with.
- We may not know whether they could reproduce the reasoning without assistance.
This does not mean Student B has learned nothing.
Nor does it mean Student A should automatically receive the higher grade.
The issue is different:
The final product may not provide enough evidence to tell us what happened during the learning process.
-----------------------------------------------------------------------------------
From One Performance to Many Pieces of Evidence
The faculty member proposed a different possibility. What if learning were demonstrated gradually?
Instead of:
One major assignment → One major grade
the assessment process might look more like:
Micro-Assessment 1
↓
Micro-Assessment 2
↓
Application
↓
Reflection
↓
Micro-Assessment 3
↓
Final demonstration
↓
Cumulative evidence of learning
The idea was not to create more tests. It was to create more meaningful evidence.
A student might first explain a concept.
- Then apply it.
- Then analyse an example.
- Then critique a solution.
- Then reflect on what they had misunderstood.
- Then demonstrate the concept in a new context.
- The individual activities could be relatively small.
But together they could reveal something a single final assignment cannot easily show:
How the student's understanding developed.
-----------------------------------------------------------------------------------
Micro-Assessment Is Not Micro-Grading
The university quickly recognised a potential problem.
If every small activity carried marks, students might simply become obsessed with many small grades instead of one large grade.
The institution could end up with:
More assessment.
without necessarily producing:
Better evidence of learning.
So the distinction became important.
A micro-assessment does not necessarily need to be heavily graded.
It could be:
- Formative - The student receives feedback but no grade.
- Diagnostic - The teacher uses the response to identify misconceptions.
- Low-stakes - The activity contributes only a small amount to the final grade.
- Reflective - The student examines how their understanding has changed.
- Evidence-building - The response becomes part of a larger portfolio of learning.
The purpose is not to make students submit something every few days. The purpose is to create purposeful opportunities to observe learning.
Formative assessment literature has long emphasised the role of evidence, feedback and classroom interaction in supporting learning rather than treating assessment simply as a final judgement.
The university therefore reframed the idea:
The goal is not more assessments. The goal is better evidence.
The Assessment Trail
Consider a student learning software architecture.
Instead of one large project at the end of the semester, the teacher creates a sequence of activities.
Activity 1 : Explain what a software architecture pattern is in your own words.
Activity 2 : Compare two architecture patterns and identify when each might be appropriate.
Activity 3 : Analyse an existing software design and identify its weaknesses.
Activity 4 : An AI tool has proposed a solution. Evaluate the solution and identify anything that should be changed.
Activity 5 : Design your own solution for a new scenario.
Activity 6 : Explain the decisions you made and defend them against an alternative approach.
Now the teacher has something different.
Not simply a final project.
A trail of evidence.
The teacher can see:
- what the student understood initially,
- how they applied the concept,
- where misconceptions appeared,
- whether they could critique a solution,
- whether they could make independent decisions,
- and whether they could explain and defend those decisions.
The final product still matters.
But it is no longer the only window into learning.
Where Does AI Fit?
This brought the university back to its original problem.
If students can use AI throughout the process, does cumulative assessment actually solve anything?
Not necessarily.
But perhaps that is the wrong question.
The better question is:
What role is AI playing in each activity?
Suppose the learning outcome is:
Students will independently explain a fundamental programming concept.
Then unrestricted AI assistance during that particular assessment may undermine what the teacher is trying to measure.
But suppose the learning outcome is:
Students will evaluate and improve computer-generated solutions.
Now AI may be part of the assessment itself.
The distinction is important.
The question is not simply:
Was AI used?
It is:
What capability was the student expected to demonstrate?
This is consistent with the broader principle of constructive alignment: assessment needs to produce evidence relevant to the intended learning outcome.
AI does not remove that principle.
It makes it harder to ignore.
The university began experimenting with assessment tasks that made student thinking more visible.
Instead of asking students only to submit a finished answer, teachers asked them to:
explain how they reached an answer,
compare alternative solutions,
critique an AI-generated response,
justify a decision,
or reflect on how their understanding had changed.
The purpose was not to catch students using AI.
It was to obtain better evidence of learning.
For example, a teacher might give students an AI-generated explanation of a historical event alongside two reliable historical sources.
Students could then be asked:
- Which claims are supported?
- What has been oversimplified or omitted?
- What would you change?
Here, AI is no longer simply producing an answer.
It becomes an object of investigation.
The assessment is now measuring something different: the student's ability to evaluate, question and make judgements.
Micro-Assessments + Cumulative Evidence
This led to the central idea of the case.
Instead of depending heavily on one major assignment, learning could be demonstrated through a series of purposeful activities:
Explain → Apply → Analyse → Reflect → Demonstrate
Each activity does not necessarily need to carry a significant grade.
Some could be formative.
Some could be low-stakes.
Some could simply provide evidence of progress.
Together, they create a cumulative picture of learning.
The goal is not:
More assessment.
It is:
Better evidence.
A student's final submission would still matter.
But it would no longer have to carry the entire burden of proving what that student has learned.
But Is This Really a Solution?
Not automatically.
More frequent assessment can increase teacher workload.
Too many small assessments can overwhelm students.
And if every activity becomes another graded task, we may simply create more opportunities to chase marks.
So the idea is not to replace one large assessment with twenty small ones.
The real question is:
What evidence do we actually need to determine whether the learning outcome has been achieved?
That evidence might be an explanation, an application, a critique, a design decision, a reflection or a demonstration.
The assessment should follow the learning objective — not the other way around.
When AI Does Not Disappear
This approach also does not eliminate AI. And perhaps it should not.
- For some learning outcomes, students may need to work without AI.
- For others, AI may be a legitimate learning tool.
- In some activities, students may even be asked to use AI and then evaluate its output.
The important questions become:
- What is the student expected to learn?
- What role is AI playing?
- What must the student demonstrate?
- What evidence will show that learning?
The question has therefore shifted from:
“Was AI used?”
to:
“Can we still see the learning?”
The Bigger Educational Management Question
The case began as an AI problem.
It ended as an assessment question.
If assessment influences student behaviour, then institutions need to ask what their assessment systems are actually encouraging.
Are students being encouraged to:
learn → practise → improve → demonstrate
or primarily to:
produce → submit → get a grade → move on?
Perhaps the answer is not to make students care less about grades.
Perhaps it is to make the learning behind the grade more visible.
That may require a shift from isolated high-stakes performances towards a more balanced combination of formative activities, meaningful assessment and cumulative evidence.
What This Case Taught Me
Assessment does more than measure learning.
It also communicates what an institution considers important.
Generative AI has made the limitations of product-focused assessment much harder to ignore.
But perhaps the bigger opportunity is to ask a question that goes beyond AI:
If we want to know what students have learned, are we collecting enough evidence to actually see it?
Maybe the goal is not to make assessment harder for students.
Maybe it is to make assessment more meaningful.
Because when the grade becomes the goal, students may learn to optimise the product.
But when learning becomes the goal, assessment can become evidence of the journey.
The question is not simply how we stop students from using AI.
The question is whether our assessment system gives us enough evidence to know that they have learned.
-----------------------------------------------------------------------------------
Research Behind This Approach
The ideas in this case are not based only on my classroom experience. They are also connected to several established ideas in educational research.
Constructive alignment is one useful starting point. John Biggs argued that teaching, learning activities and assessment should be aligned with what students are actually expected to learn. In other words, assessment should not simply measure the easiest thing to produce; it should provide evidence of the intended learning.
This becomes particularly important when the intended learning involves analysing, applying, evaluating, creating or making decisions. A polished final product may look impressive, but if the assessment does not allow us to see the thinking behind it, we may have limited evidence of whether that learning actually occurred.
A second important idea comes from the work of Paul Black and Dylan Wiliam on formative assessment. Their work places assessment within the learning process rather than treating it only as something that happens after learning. Formative assessment can provide information that helps teachers and students understand where learning is, what needs attention, and what should happen next.
This supports the principle behind Micro-Assessments + Cumulative Evidence used in this case.
The purpose is not to create more examinations or turn every classroom activity into a graded event. It is to create several meaningful opportunities to see learning developing over time.
A student might first explain an idea, then apply it, then analyse a problem, receive feedback, revise their thinking, and finally demonstrate what they can do independently. Each piece provides some evidence. Together, they can provide a richer picture of learning than a single final performance.
The arrival of generative AI has made this question even more important.
Lye and Lim (2024), writing specifically about assessment in tertiary education, argue that simply trying to police generative AI or engage in an “arms race” with AI-detection tools does not address the deeper assessment problem. They propose reconsidering assessment design according to the purpose of the learning and the appropriate role of AI. Their framework includes assessments where AI is prohibited, assessments designed around tasks AI currently handles poorly, and assessments that deliberately incorporate AI as part of the learning activity.
This reinforces an important distinction in this case:
The question is not simply whether AI was used. The question is whether the assessment still gives us meaningful evidence of the student's learning.
UNESCO's guidance on generative AI in education similarly emphasises a human-centred approach, including meaningful pedagogical use, ethical considerations, and the need for educational institutions to develop appropriate policies and practices rather than treating AI simply as a technological problem. UNESCO also highlights the need to reconsider how knowledge, learning and assessment are understood in the age of generative AI.
None of this means that Micro-Assessments + Cumulative Evidence is a universal solution. More assessment can also create more workload. Poorly designed micro-assessments can simply become a collection of small grades, increasing pressure without improving learning.
The important principle is therefore not more assessment.
It is better evidence of learning.
The research, together with my own classroom experience, leads me back to the question at the heart of this case:
If we want to know what students have learned, are we collecting enough evidence to actually see it?
Perhaps assessment should not try to capture the whole journey in one final product.
Perhaps it should help us see the journey while it is still happening.
References
Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32, 347–364.
Read the article
Black, P., & Wiliam, D. (2009). Developing the theory of formative assessment. Educational Assessment, Evaluation and Accountability, 21, 5–31.
Read the article
Lye, C. Y., & Lim, L. (2024). Generative Artificial Intelligence in Tertiary Education: Assessment Redesign Principles and Considerations. Education Sciences, 14(6), 569.
Read the open-access article
Miao, F., & Holmes, W. (2023). Guidance for Generative AI in Education and Research. UNESCO.
Read the UNESCO guidance
What do you think? 👇
I'd be interested to hear from teachers, educators, lecturers, academic leaders and other education professionals who have experience with these issues. If you have a different perspective, a classroom experience, or an approach that has worked for you, please feel free to share it in the comments.
Thoughtful disagreement is welcome too. The aim is to learn from one another and keep the conversation going. Thank you!

Comments
Post a Comment