Pearson Accelerator

Writing Coach

Context

On Pearson’s Accelerator team, which takes AI concepts from problem definition to tested prototypes that shape what the business builds, I designed an AI writing coach for PTE Academic learners.

Writing is the hardest part of the exam to improve alone. A reading answer is right or wrong, but an essay comes back as a number with no visible reason attached. The feedback that would actually teach you, why an argument is thin or why a sentence undercuts your point, is exactly what a busy learner rarely gets in time. The obvious fix, handing the essay to an LLM, has a subtler failure: general-purpose models default to correcting sentences and handing back a tidier version, which quietly does the learner’s thinking for them.

This proof of concept explored whether an AI coach could support the whole act of writing, from finding a position to drafting, revising and understanding a score, in a way learners would trust and, more importantly, learn from.

The Writing Coach brainstorm screen, with the essay prompt and the conversational AI

Problem

Learners preparing for PTE Academic can practise writing endlessly and still not improve, because the thing that teaches, specific and trustworthy and timely feedback, is the thing they can’t get.

Feedback is scarce and slow

No. 01

A score tells you where you landed, not what to change. Real feedback on an essay depends on a teacher’s time, which most learners don’t have on demand.

AI corrects, it doesn’t coach

No. 02

General-purpose LLMs default to sentence-level correction and position the learner as a novice to be fixed. A cleaner essay comes back, but the skill doesn’t transfer.

The blank page

No. 03

Before any of that, learners stall at the start, unsure how to take a position or structure an argument for the prompt in front of them.

How might we give learners writing feedback they trust and learn from, coaching the whole process, without quietly writing the essay for them?

Grounding the AI in research

The feedback design is not a hunch. It is built on a Stanford study that pinpoints exactly where AI writing feedback goes wrong. Mah, Tan, Phalen, Sparks and Demszky compared how large language models and expert teachers give feedback on student writing, using a framework of dialogic feedback across cognitive, social and structural dimensions.

Their finding became the hinge of this project. LLMs mostly enact corrective feedback at the sentence level and position students as novices to be remediated. Teachers, by contrast, give feedback at multiple levels and position students as agentic writers who can revise their own work. In other words, the default behaviour of an LLM is the opposite of what actually teaches someone to write.

So I set out to make the coach behave less like the model and more like the teacher: work at the level of argument before grammar, hand the thinking back to the learner, and never simply supply the corrected sentence.

“LLMs primarily enacted corrective feedback at the sentence level and positioned students as novices requiring remediation. Teachers positioned students as agentic writers.”

Mah, Tan, Phalen, Sparks & Demszky, Stanford University. EdWorkingPaper No. 25-1193

Market context

I audited the writing-help tools learners actually reach for, and annotated each flow. Grammarly corrects mechanics but knows nothing about the PTE rubric or the exam a learner is working toward. Khan Academy’s Khanmigo coaches more conversationally, but isn’t built around a specific test. PTE-specific tools score essays against the rubric, yet tend to hand back a corrected version rather than teach the learner to produce it.

One pattern ran through all of them: they either correct or score, and the more they do for the learner in the moment, the less the learner is left able to do alone. Almost none coach across the whole process, from finding a position to revising a draft, while deliberately keeping the pen in the learner’s hand. Generation and correction are the commodity. Feedback that builds a writer who can perform without the tool is not.

Process

One coach, two connected phases

Rather than bolt feedback onto a finished essay, I designed a flow that follows the writer from a blank page to a graded, understood result: Brainstorm, Write, Review, Score. It runs in two connected phases, first understand the topic, then write with that understanding, and every AI moment follows one rule: help the learner do the work, never do it for them.

Phase one, Brainstorm: a conversational stage comes first, because a weak PTE essay is usually a thin argument, not just shaky grammar. Scaffolds push the learner to take and defend a position rather than be handed one, and they can work in English, Spanish, Portuguese, Hindi or their own language. They leave with a position they understand, not a paragraph the AI wrote.

Phase two, Write: the same coach carries the brainstorm into the editor. It already knows the position taken and the arguments weighed, so support during drafting is specific to their essay rather than a fresh chatbot that has never met them.

Review: feedback deliberately withholds the answer. A flagged sentence tells the learner that something is wrong and helps them see why, but they work out the fix and make the edit themselves. If they want to go further, they can practise the underlying rule with a short quiz instead of accepting a correction.

Score: the essay is scored against the real PTE rubric, each category legible rather than one opaque mark, with a before/after card that shows the score moving as the learner revises.

The full flow: brainstorm, the essay editor, the sentence-level review, and the scored result
The flow end to end: brainstorm, write, the sentence-level review, and the scored result

Designing against the model, on purpose

Every one of those decisions traces back to the Stanford finding. The review stage never hands over the answer: it points the learner at the problem, makes them fix it, and offers a quiz to drill the rule behind it. The brainstorm builds real understanding before a word is written. Together they coach the way a teacher would, and the AI’s job is to provoke and support the thinking, never to replace it.

User testing

I ran an unmoderated study on UserTesting.com with six participants, all non-native English speakers preparing for an English proficiency exam, the actual audience rather than a convenient proxy. The shared essay was seeded with deliberate errors so the review and scoring flow had real material to catch, so finding and fixing them counted as success, not a bug report. Each session ran the full flow, landing to final score, and closed with a Likert battery on ease, confidence, accuracy and usefulness.

The voice toggle hid

No. 01

Four of six never noticed they could speak; one found it unprompted and loved it. A discoverability problem, not a value problem.

‘I thought it would write the essay’

No. 02

Two participants expected the AI to produce the essay outright; one acted on it and went off-prompt. A sign the coach’s role needed anchoring before the open chat invited the wrong request.

Trusted on sight

No. 03

The red-flag convention needed no explanation, and ‘explain in my language’ was singled out as valuable by the ESL testers: the audience-specific bet paying off.

Agency, validated

No. 04

Given the choice, the majority chose to fix the flagged sentence themselves rather than be shown. The dialogic wager held up in behaviour, not just theory.

Did the score update?

No. 05

Two testers weren’t sure the score had recalculated after their edits, denting trust in the before/after card. A ‘Rechecked’ confirmation resolves it.

I ranked every issue by impact against effort into a prioritised backlog, and kept the disagreements rather than averaging them away. One strong tester didn’t want AI help before a ‘real’ submission, while others used it proactively: a genuine split between a test-condition and a learning-tool mindset that is worth designing for rather than smoothing over.

Iteration: what testing changed

The six sessions produced a prioritised set of changes. Four shaped the next version most, each traced back to a finding.

Voice, made impossible to miss. Four of six never noticed they could talk to the coach; the option lived in an unlabelled grey toggle that read as decoration. I redesigned it into an icon-led control, a keyboard glyph for Type and a magenta microphone for Speak, and moved it into the input’s line of sight so it announces itself. Paired with the ‘Speak’ label, the mic reads as a mode, not a request to record a voice note.

A landing page for context. Learners arrived unsure what the tool actually was. I added a landing page that frames what the coach does and does not do, so people understand they are about to be coached, not handed an essay, before they begin.

An opening screen that emphasises the prompt. Two participants expected the AI to write the essay outright. So the flow now opens on a focused screen showing only the essay prompt and a single input, putting the question front and centre before the open chat can invite the wrong request.

A score that explains itself. The score first read as one opaque number, and some testers weren’t sure what it meant or whether it had changed. I added a score reveal that breaks the mark down by rubric category with a short explanation of each, so the number becomes something the learner can act on.

The changes in place: the landing page, the redesigned toggle with a labelled microphone, and the score reveal that breaks the mark down by rubric category
The changes in place: the landing page, the redesigned voice toggle, and the score that explains itself

Final design

A high-fidelity, interactive prototype of a writing coach that stays with the learner across the whole task: a brainstorming partner that helps them find a position without drafting for them, a calm editor with the AI on hand, a review stage that flags and explains and offers practice rather than just corrections, and a rubric-based score with a before/after card that makes progress visible.

The full flow: brainstorm, write, review, score, and the before/after card

Outcome and impact

For an Accelerator proof of concept, success was a de-risked decision, and the study delivered one. Honest limits first: six participants, unmoderated, one likely rendering bug still to QA, and no live usage metrics. This is a validated concept, not a shipped product.

The core bet, validated

Learners wanted coaching across the process and said they’d use it regularly; the AI panel was rated helpful, not distracting, by nearly everyone.

An AI point of view

The feedback was designed against a documented LLM limitation by withholding the answer and making the learner fix their own error, with a quiz to drill the rule.

A study with teeth

Seeded errors, severity-rated findings, a prioritised impact and effort backlog, and preserved disagreement rather than a flattened average.

ESL-first, and it held

The choices made for non-native speakers, multilingual brainstorming and ‘explain in my language’, were the ones testers singled out as valuable.

What I would do differently

Anchor the coach’s role before the conversation opens. The ‘it’ll write my essay’ misconception was predictable; a one-line statement of what the coach does would have prevented the off-prompt session entirely.

Test the trust question with a moderated round. The unmoderated format is efficient for usability but thin on the why behind trust; one moderated session would have told me more about the test-condition versus learning-tool split than any rating.

Test the brainstorm earliest. It is the stage furthest from a correct answer and the easiest to get subtly wrong, so it deserved the first rough prototype rather than a share of a polished one.