Getting every student substantive feedback on a draft when you have ninety of them

English English as a Second Language Intermediate 90 min Free. The PAIRR prompts are openly licensed and work on any free tier.

The situation

By the time I have written real comments on eighty drafts the revision deadline has passed and my feedback is an autopsy rather than a coaching note. What I actually want to assess is not whether the draft is good — it is whether the writer can read feedback, weigh it, and make a defensible revision decision. This moves my reading time from the draft to the reflection.

Steps

  1. Set the policy and the readings before any draft exists

    LMS, short readings on hallucination and bias

    Spend one session on your AI policy plus readings on limitations, bias and environmental cost, and have students write a brief reflection on how they intend to use AI in the course. This is load-bearing, not throat-clearing.

    What you only learn by doing it: Do this before they have a draft in hand. Once something is due, the readings read as a hoop. The reflection also gives you a baseline writing sample in their own voice, which is quietly useful later.

  2. Run peer review first — never after the AI

    Canvas peer review, Google Docs, or in-class exchange

    Each student gives and receives feedback from two peers before any AI touches the draft. Students consistently describe peer feedback as the deeper reader, because peers know the assignment, the course and the classroom context the model has no access to.

    What you only learn by doing it: This ordering is the single most important thing in the workflow. If AI goes first students anchor on it and peer review collapses into agreement with the machine. Colleagues who reversed the order to save class time watched peer comments go from paragraphs to “I agree with what the AI said.”

  3. Give students a rubric-loaded prompt, not a blank chatbot

    Any free tier, plus your actual rubric

    Students paste a structured prompt containing your rubric and their draft. The prompt frames comments as “a reader might wonder…” rather than as rules, instructs the model not to supply language students can paste in, and caps the response at roughly two strengths and two considerations tied to quoted passages.

    What you only learn by doing it: Output quality tracks your rubric far more than your prompt. Around 31% of student reflections coded the AI feedback as generic — and generic feedback is almost always downstream of a generic rubric. If the response is vague, rewrite the criteria, not the prompt.

  4. Require a comparison reflection that names disagreements

    LMS assignment or a form

    Students compare peer and AI feedback: where they agree, where they conflict, which advice they are rejecting and why. This is the assessed artifact — it is where you can see whether a student can evaluate advice rather than obey it.

    What you only learn by doing it: Ask explicitly for at least one piece of AI advice they are refusing, with a reason. Without it most students submit compliance narratives. But grade the reason, not the presence of an objection — the study's own authors note that asking about disagreement may manufacture the appearance of critical engagement.

  5. Read reflections; comment on drafts selectively

    Your gradebook and a short list of flagged students

    Read every reflection; comment fully only where the reflection shows a student is stuck, misreading their own writing, or has accepted bad advice. This is where the time comes back.

    What you only learn by doing it: Budget your reclaimed time toward the students whose reflection is three sentences long. Those are the ones who read neither source. The reflection is a better triage signal than draft quality.

Where this breaks down

The clearest negative finding in the research is that AI feedback was only rated as useful as peer feedback in the presence of human feedback — only 6% of students preferred AI alone, and the authors are direct that removing humans from the loop contradicts what we know about how relationships drive engagement.

The linguistic-justice risk is real and under-measured. Models default to Standard Academic English and will quietly push multilingual writers toward it. The PAIRR authors say plainly that they did not analyse their outputs for bias and did not warn students — so you should.

Students will use AI on the reflection itself. The study found three suspected cases, and the right response is assessment design — in-class writing, a short conference, specificity requirements generic text cannot satisfy — not detection.

A rubric written for human readers often makes a poor prompt. If your criteria say “engages meaningfully with sources,” expect meaningless feedback back.

Provenance: documented, peer-reviewed practice. Peer and AI Review + Reflection (PAIRR), MacArthur, Minnillo, Sperber, Stillman & Whithaus, Computers and Composition — 654 students across 10 writing courses, 68% multilingual. Community college implementation by Anna Mills at College of Marin.