Pearson Accelerator
Writing Coach
Context
On Pearson’s Accelerator team, which takes AI concepts from problem definition to tested prototypes that shape what the business builds, I designed an AI writing coach for PTE Academic learners.
Writing is the hardest part of the exam to improve alone. A reading answer is right or wrong, but an essay comes back as a number with no visible reason attached. The feedback that would actually teach you, why an argument is thin or why a sentence undercuts your point, is exactly what a busy learner rarely gets in time. The obvious fix, handing the essay to an LLM, has a subtler failure: general-purpose models default to correcting sentences and handing back a tidier version, which quietly does the learner’s thinking for them.
This proof of concept explored whether an AI coach could support the whole act of writing, from finding a position to drafting, revising and understanding a score, in a way learners would trust and, more importantly, learn from.
Problem
Learners preparing for PTE Academic can practise writing endlessly and still not improve, because the thing that teaches, specific and trustworthy and timely feedback, is the thing they can’t get.
Feedback is scarce and slow
No. 01A score tells you where you landed, not what to change. Real feedback on an essay depends on a teacher’s time, which most learners don’t have on demand.
AI corrects, it doesn’t coach
No. 02General-purpose LLMs default to sentence-level correction and position the learner as a novice to be fixed (Mah et al., Stanford, 2025). A cleaner essay comes back, but the skill doesn’t transfer.
The blank page
No. 03Before any of that, learners stall at the start, unsure how to take a position or structure an argument for the prompt in front of them.
How might we give learners writing feedback they trust and learn from, coaching the whole process, without quietly writing the essay for them?
Market context
The writing-help market splits into two camps that both miss this. Grammar tools like Grammarly correct mechanics but know nothing about the PTE rubric or the exam a learner is working toward. Exam-prep and general assistants can produce or polish an essay, but they optimise for the output, not the writer: the more helpful they are in the moment, the less the learner is left able to do alone.
Almost none coach across the whole process, from brainstorm to draft to revision to understanding the score, while deliberately preserving the learner’s agency at each step. Generation and correction are the commodity. Feedback that builds a writer who can perform without the tool is not.
Process
One coach, two connected phases
Rather than bolt feedback onto a finished essay, I designed a flow that follows the writer from a blank page to a graded, understood result: Brainstorm, Write, Review, Score. It runs in two connected phases, first understand the topic, then write with that understanding, and every AI moment follows one rule: help the learner do the work, never do it for them.
Phase one, Brainstorm: a conversational stage comes first, because a weak PTE essay is usually a thin argument, not just shaky grammar. Scaffolds push the learner to take and defend a position rather than be handed one, and they can work in English, Spanish, Portuguese, Hindi or their own language. They leave with a position they understand, not a paragraph the AI wrote.
Phase two, Write: the same coach carries the brainstorm into the editor. It already knows the position taken and the arguments weighed, so support during drafting is specific to their essay rather than a fresh chatbot that has never met them.
Review: feedback deliberately withholds the answer. A flagged sentence tells the learner that something is wrong and helps them see why, but they work out the fix and make the edit themselves. If they want to go further, they can practise the underlying rule with a short quiz instead of accepting a correction.
Score: the essay is scored against the real PTE rubric, each category legible rather than one opaque mark, with a before/after card that shows the score moving as the learner revises.
Designing against the model, on purpose
The through-line is a Stanford paper (Mah et al., 2025). Its finding, that LLMs correct at the sentence level and position students as novices while teachers work dialogically and position them as agentic writers, became the brief for the feedback design specifically. The review stage never hands over the answer: it points the learner at the problem, makes them fix it, and offers a quiz to drill the rule behind it. Paired with a brainstorm that builds real understanding before a word is written, the product coaches the way a teacher would, and the AI’s job is to provoke and support the thinking, never to replace it.
User testing
I ran an unmoderated study on UserTesting.com with six participants, all non-native English speakers preparing for an English proficiency exam: the actual audience, not a convenient proxy. The shared essay was seeded with deliberate errors so the review and scoring flow had real material to catch, and finding and fixing them counted as success. The study covered the whole flow, landing to final score, and closed with a Likert battery on ease, confidence, accuracy and usefulness.
The voice toggle hid
No. 01Four of six never noticed they could speak; one found it unprompted and loved it. A discoverability problem, not a value problem.
‘I thought it would write the essay’
No. 02Two participants expected the AI to produce the essay outright; one acted on it and went off-prompt. A sign the coach’s role needed anchoring before the open chat invited the wrong request.
Trusted on sight
No. 03The red-flag convention needed no explanation, and ‘explain in my language’ was singled out as valuable by the ESL testers: the audience-specific bet paying off.
Agency, validated
No. 04Given the choice, the majority chose to fix the flagged sentence themselves rather than be shown. The dialogic wager held up in behaviour, not just theory, and a real split surfaced between a test-condition and a learning-tool mindset worth designing for.
Iteration
The study produced a prioritised set of changes, ranked by impact against effort. The high-impact, low-effort fixes:
Label the voice option: enlarge it and say “Speak instead,” turning a hidden icon into an offer, and move it into the input’s line of sight rather than a top corner.
Anchor the prompt and the coach’s role: pin the essay prompt above the chat and have the coach restate what it will and won’t do, closing the ‘it’ll write it for me’ gap.
Auto-explain on click: one click on a flagged sentence should explain it, not two.
Confirm the recheck: a “Rechecked” state so the before/after score is visibly earned, protecting trust in the product’s best feature.
Final design
A high-fidelity, interactive prototype of a writing coach that stays with the learner across the whole task: a brainstorming partner that helps them find a position without drafting for them, a calm editor with the AI on hand, a review stage that flags and explains and offers practice rather than just corrections, and a rubric-based score with a before/after card that makes progress visible.
Outcome and impact
For an Accelerator proof of concept, success was a de-risked decision, and the study delivered one. Honest limits first: six participants, unmoderated, one likely rendering bug still to QA, and no live usage metrics. This is a validated concept, not a shipped product.
The core bet, validated
Learners wanted coaching across the process and said they’d use it regularly; the AI panel was rated helpful, not distracting, by nearly everyone.
An AI point of view
The feedback was designed against a documented LLM limitation by withholding the answer and making the learner fix their own error, with a quiz to drill the rule.
A study with teeth
Seeded errors, severity-rated findings, a prioritised impact and effort backlog, and preserved disagreement rather than a flattened average.
ESL-first, and it held
The choices made for non-native speakers, multilingual brainstorming and ‘explain in my language’, were the ones testers singled out as valuable.
What I would do differently
Anchor the coach’s role before the conversation opens. The ‘it’ll write my essay’ misconception was predictable; a one-line statement of what the coach does would have prevented the off-prompt session entirely.
Test the trust question with a moderated round. The unmoderated format is efficient for usability but thin on the why behind trust; one moderated session would have told me more about the test-condition versus learning-tool split than any rating.
Test the brainstorm earliest. It is the stage furthest from a correct answer and the easiest to get subtly wrong, so it deserved the first rough prototype rather than a share of a polished one.