Four Months in the Lab

Four Months in the Lab

One of GAILE’s primary goals: to discover the role generative AI could play a role in the curriculum; in assessment, in the student learning experience, in the review, and everywhere in between.

In April we crafted our methodology, and an argument that a lab like ours should produce defensible insight rather than just good stories. It assumed two things: that we could co-design with educators, and that we could collect clean signal while doing it. But, 4-ish months of this endeavor has been truly eye opening, and the methodology in retrospect feels a bit more like ideology.

This piece is about what we've learned so far.

1. Co-creation is great, but hard to scale

In May, we invited the community of educators to a co-creation workshop where we endeavored to identify the problems we'd like most to be solved, and what solutions to them might look like. Despite the excitement around the event educators, the ones we most want to empathise and co-design with, are busy teaching.

The people who did show up were generous and surfaced grounded insights: academics 'need their own personal learning designer,' differentiated design is unaffordable, students are optimising for output while reasoning goes invisible is offloaded to AI, work feels "completely written by AI...they're just completely dropping the learning."

If you're reading this and you're close to these problems, they'll feel obvious. But elucidating and communicating these is core to the lab's work. If one of us leaves, there's a trail of stories to rebuild empathy from, paramount for personalised design and anti AI-slop. And, if we make unconventional decisions, we can point to literal evidence.

2. We found no signals for using transcriptions in assessments...yet

High-quality transcriptions are now easier than ever to produce. Even better, the infrastructure exists at RMIT. In April we set out to find how educators might use them in their grading workflows for viva style assessments, which, are great (not perfect of course) at supporting “closed assessment” and assuring learning. Unfortunately, they're very hard to scale.

It was immediately obvious after talking to educators: actually watching submitted recordings is paramount. Subtitles for recordings might be nice, but it likely wouldn't be a supplement.

We paused this. AI isn't likely to help here yet, and it may never be. How could we calibrate trust in this kind of AI-enabled experience? Could we ever be confident about giving someone a HD purely based on the outputs of an AI, even if they're transcriptions? What are educators asking themselves, from the video, that isn’t yet revealed in text-only transcript?

3. Chat isn't always the right container

A secret GAILE mantra this year was “surely AI is more than just chat bots”. So in collaborating with Anselm Paul (Senior Manager, Learning Design and Development in STEM), we explored the idea of scaling a VAL persona (RMIT’s in house AI Platform) called Assessment Authentifire: a chatbot able to generate standardised drafts for authentic assessments. The small community of educators and program managers enjoyed the suggestions for authenticity and the back and forth refining assessments, treating early drafts as "concrete starting points" or "quality first drafts."

However, the user experience of chat for these particular tasks often worked against them. They reported overwhelming "walls of text", that their personalised nuances, contexts, and conventions weren't carried between assessments which meant calibration had to be constantly repeated. Iteration was effortful rather than incremental; several educators needed multiple attempts just to land small changes. In one case, hallucinated references surfaced "despite multiple iterations". Generating several drafts in the same thread produced "significant overlap" rather than progressively better work, and the memory mechanism behind each chat was less than ideal.

So, we put authentifire in the lab and toyed around with the most immediately obvious designs, in pursuit of addressing feedback. What if: we put those same authentifire instructions, i.e. the recipe for what a good assessment looks like at RMIT, into a different vessel? A document editor or a multi-step form? 

Split-screen design mock-up showing a 'Document-first editor' on the left and a 'Step-by-step wizard' on the right. The editor displays an uploaded course guide and a 'Problem Solving Workshop' assessment divided into sections. The wizard is at step 3 of 6, 'Define the Assessment Task', with fields for the real-world problem, the student’s professional role, an artefact or performance and why the task matters.Split-screen design mock-up showing a 'Document-first editor' on the left and a 'Step-by-step wizard' on the right. The editor displays an uploaded course guide and a 'Problem Solving Workshop' assessment divided into sections. The wizard is at step 3 of 6, 'Define the Assessment Task', with fields for the real-world problem, the student’s professional role, an artefact or performance and why the task matters.

We realised we'd be jumping way ahead too quickly. We're too small of a team to re-implement something as robust as a word processor like Microsoft Word. We at least needed stronger signals before sinking significant time into anything.

Another quirk, which our design phase picked up, is that educators create assessment artefacts and specifications in a range of diverse ways, and, we know that good design means respecting nuances. For example, Overleaf is a popular tool in STEM for crafting scientific documents, including assessments, with the full power of LaTeX and the pandoc ecosystem. We realised that clearly more research was needed to find the just-right experience sitting between VAL and the canonical tools that educators can’t live without.

Does the community want AI-generated artefacts in assessment at all? What does “good” look like for everyone at RMIT, not just STEM? Which system prompts, inputs, and user experiences make significant impact? The situation warranted a stripped back prototype to ask these foundational questions, and, crowdsource feedback.

The result is now two prototypes being piloted across learning and teaching staff at RMIT:

Assessment Reviewer

  • Find any course from RMIT’s handbook, and automatically fetch the contents
  • Pick an assessment, and automatically source it’s Canvas course content
  • Use generative AI to critically review the assessment, based on original principals in Anselm’s authentifire
  • Highlight any aspect of reviews to share where the AI got it wrong, or right

Course Guide Reviewer

Screenshot of the GAILE Lab Experiments page. The 'Assessment Reviewer' experiment is selected beside 'Course Guide Reviewer'. It is described as a framework-grounded review of a single assessment task, checked against the live Canvas course with each point linked to evidence. A 'Choose Course' button and 'How does this work?' link are visible.Screenshot of the GAILE Lab Experiments page. The 'Assessment Reviewer' experiment is selected beside 'Course Guide Reviewer'. It is described as a framework-grounded review of a single assessment task, checked against the live Canvas course with each point linked to evidence. A 'Choose Course' button and 'How does this work?' link are visible.

What the lab environment brings us

The lab has its own lite hypotheses:

  • the ability for participants to give inline-feedback in their own time, like we comment on Word documents, will engage and respect time-poor staff (screenshot below)
  • prototypes that are explicitly called out as "going away in X weeks” to participants will change what people are willing to say, because reacting honestly to something temporary is far easier than something released without consultation
  • valuable analytics on how people are using experiments, combined with feedback, will help us make better design choices
Screenshot of an assessment review interface. It shows an overall score of 2.9 out of 5 with 40% weighting. Under 'Authentic Assessment', 'Challenges cognitively' is rated 5 out of 5, marked 'Good', with feedback explaining that the task requires higher-order thinking, ethical reasoning and systems thinking. 'Encourages collaboration' is rated 1 out of 5 and marked 'Needs-Work'.Screenshot of an assessment review interface. It shows an overall score of 2.9 out of 5 with 40% weighting. Under 'Authentic Assessment', 'Challenges cognitively' is rated 5 out of 5, marked 'Good', with feedback explaining that the task requires higher-order thinking, ethical reasoning and systems thinking. 'Encourages collaboration' is rated 1 out of 5 and marked 'Needs-Work'.

There are early signs of these working positively, with rich feedback coming in across our initial pilots. We’re seeing point by point disputes, decrees of "overreach", push-back on a request for industry relevance in a theoretical course, and discrediting Blooms in favor of other taxonomies. These are exactly the types of signals we were waiting for.

What’s next?

“I wonder what you could also do with ___ data” is a common thought voiced from some of our participants. It feels like there’s so much we could do with Canvas course content in particular, so we’re exploring similar automated reviewers that ask other kinds of questions. Personally, working on some of our education innovation projects, such as designing an AI patient simulator prototype in SHBS, and exploring the efficacy of using AI for photography judgements, has been highly elucidating. I’ll be thrilled to read and report on their pilot results.

Of this I'm certain: building empathy is how we’ll find the right way forward; the right designs; the right interventions. If you have a story to tell us, we’re always open-minded and willing to listen, and you might just unlock something we’ve been unable to crack.

Drop us a line at gaile@rmit.edu.au

Looking forward to the next reflection!

Definitions

  • GAILE: Generative AI Lab for Education
  • VAL: RMIT's in-house AI platform (Virtual Assistant for Learning)
  • SHBS: School of Health and Biomedical Sciences
  • HD: High Distinction, RMIT's top grade band
  • Bloom's taxonomy: a hierarchy of cognitive skills (remember, understand, apply, analyse, evaluate, create) commonly used to frame learning outcomes and assessment verbs

References and readings

03 September 2026

More GAILE blogs

aboriginal flag float-starttorres strait flag float-start

Acknowledgement of Country

RMIT University acknowledges the people of the Woi wurrung and Boon wurrung language groups of the eastern Kulin Nation on whose unceded lands we conduct the business of the University. RMIT University respectfully acknowledges their Ancestors and Elders, past and present. RMIT also acknowledges the Traditional Custodians and their Ancestors of the lands and waters across Australia where we conduct our business - Artwork 'Sentient' by Hollie Johnson, Gunaikurnai and Monero Ngarigo.

Learn more about our commitment to Aboriginal and Torres Strait Islander peoples