Pearson Accelerator
Teacher Assistant
Context
On Pearson’s Accelerator team, which takes cutting-edge AI concepts from problem definition to tested prototypes that shape what the business builds, I designed and prototyped an AI assistant for English language teachers.
Teachers lose roughly a working day every week to lesson preparation, and the tools promising to fix that mostly hand back material they then have to check line by line, which moves the work rather than removing it. This proof of concept explored whether generated practice could be targeted and trustworthy enough to genuinely save that time.
The project was paused before user testing, so this case study is explicit about what is evidenced and what is not.
Concept prototype · paused before validation
Problem
Teachers need a constant supply of fresh, targeted practice, and building it by hand is slow, repetitive work that eats into teaching time. AI was supposed to solve this. Mostly it relocated the effort.
Time-consuming
No. 01Teachers spend an average of 7.4 hours a week preparing lessons (OECD TALIS 2024), close to a full working day.
Repetitive
No. 02The same task types get rebuilt from scratch, class after class, with little reuse between them.
Trust costs time
No. 03Around 78% of educators check AI-generated content before using it. Generation is close to solved. Verification is what still costs teachers their evening.
How might we generate practice a teacher can trust, quickly enough that checking it doesn’t eat the time it saved?
Market context
The AI-for-teachers market is crowded and largely free. Broad suites like MagicSchool bundle dozens of generators, while specialists do one job well: Diffit adapts reading passages to level, Curipod builds interactive classroom slides, Twee focuses on English language teaching.
What almost none of them do is close the loop. They produce a standalone artifact, and know nothing about the specific students in the room or the exam those students are working toward. Generation is the commodity. Generation anchored to a level framework on one side and real learner data on the other is not.
Process
Generate, then compose
I started at the moment a teacher realises they need something for tomorrow, and worked backwards to what would have to be true for generated material to actually get used rather than quietly rewritten. The design focused on three main areas:
Three surfaces, not one: a formal exam-aligned test, a quick low-stakes quiz, and an open canvas for bespoke composition. Most competitors force every job through a single generation model.
The teacher stays in control: every student-facing action is initiated or approved by the teacher. The assistant proposes, the teacher decides.
Editable by default: generated content is a starting point, not a take-it-or-leave-it output. If a teacher cannot quickly bend it to their class, the verification problem simply reappears somewhere else.



The three share one engine and differ only in how much structure they impose. The test generator constrains the most because exam alignment is the point. The canvas constrains the least because bespoke composition is the point. Forcing either job through the other surface produces something worse than both.
The quiz generator was an existing product designed by a colleague. My work there was integration rather than authorship: bringing it under one navigation, one visual language and one shared generation engine, so a teacher moves between quick quiz, formal test and open canvas without feeling handed off between three different tools. Deciding what already exists should be absorbed rather than rebuilt is part of the design work.
Closing the loop
The market gap I identified was that generators know nothing about the students in the room. So generation sits next to a class diagnostic: average PTE skill mastery across the class, broken down to the level material is actually generated at. A teacher can see that listening inferencing is the weak band and generate straight against it, rather than guessing what to practise next.

What the learner-side pilot taught me
Shortly before this I designed and tested an AI reading tutor for learners on the same team, and carried two lessons straight into this concept. The first was discoverability: I had put the AI behind a small floating control and testers repeatedly failed to notice it existed, so here generation is not hidden behind an assistant you have to find, it is the surface itself. The second was trust: learners valued AI feedback when it showed its reasoning and pointed to the evidence, which is why editability and teacher approval are structural here rather than cosmetic.
This is reasoning carried across projects, not evidence about teachers. Teachers were never tested.
Validation plan
The concept was paused before any research ran, so there are no findings. What exists is the study I designed: a five-participant unmoderated study with language teachers, built around where lesson prep actually hurts, whether teachers would want this, whether they could use it, and how they feel about AI producing material for their students.
Two decisions matter most. I quota’d for AI sceptics, requiring at least two participants who rarely or never use AI, because recruiting only enthusiasts would have produced a comfortable and useless answer. And I set the kill condition in advance: if the median trust rating came back at 3 or below out of 5, then trust, not usability, was the problem to solve before any build.
Final design
A high-fidelity, interactive prototype of a teacher’s workspace: an exam-aligned test generator working in score bands and real exam item types, a quick quiz generator, and an open canvas where AI generates composable activities that the teacher rearranges and edits.
Class and student management sit alongside, so material can point at particular learners rather than a generic average. I designed and built the prototype end to end, with the exception of the quiz generator, an existing product by a colleague that I brought into the workspace.
Outcome and impact
This concept was paused before user testing, so I cannot claim validated outcomes for it. What I can point to is that the problem it framed proved real: Pearson has since publicly launched Smart Lesson Generator, an AI tool that generates supplemental activities for English language teachers, aligned to Global Scale of English objectives. I make no claim that this work produced that product.
It does tell me the bet was worth making, and it sharpens what I would test today: not whether teachers want generated material, which the market has answered, but whether they trust it enough to hand to students without checking it line by line.
A concept made tangible
An interactive prototype that turned an abstract strategy into something the business could react to.
A point of view on AI
Teacher approval, editable output and generation as the surface itself, derived from what the learner-side pilot revealed.
A validation plan with teeth
A study designed with an AI-sceptic quota and a pre-agreed condition that would have killed the concept.
The problem proved real
Pearson has since shipped publicly in adjacent territory, confirming the problem was worth the bet.
What I would do differently
Three things, in order of how much they would have changed the work.
Start with teachers, not with the product: I framed this from secondary evidence and domain knowledge rather than first-hand conversations. Even three interviews about how teachers actually prepare materials would have grounded every decision that followed, and would have told me which of the three surfaces mattered most.
Test the trust question first, not the generation question: I designed the generation flows first because they were the visible product. The evidence says verification is where teachers lose their time, so trust was always the riskiest assumption and should have been attacked first.
Put something rough in front of people sooner: the concept reached high fidelity before it met a single user. A sketch in week one would have been worth more than a polished prototype shown never.




