Grade the specification, not the program: prompt-first CS1 with a refutation round
The situation
My students can get working code out of an assistant for anything I assign, and the honest response is not to ban it — it is to move the graded artifact upstream to the thing they are actually bad at, which is saying precisely what they want. The data on the second half is sobering: when LLM-generated code contains a bug, CS1 students correctly predict its behaviour about 15% of the time.
Steps
-
Present the problem as an image with no text
A worked input/output diagram, a picture of the desired transformation, a sketch of the data — deliberately with no textual description. Students cannot paste the problem into the model; they have to construct the description themselves.
What you only learn by doing it: The no-text rule is the entire mechanism and it collapses the moment you add a caption. If students genuinely cannot read the image, the image is the problem — redraw it, and pilot every diagram on one student first.
-
Have students write and submit the spec before any generation
LMS text submission with a sentence starter
The prompt is submitted and graded as a specification: does it state inputs, outputs, edge cases and constraints unambiguously? Grade it before they know whether it worked.
What you only learn by doing it: Cap the word count — 60 words works. Students in published pilots wrote 101-word prompts where 25 sufficed, and verbosity is not precision; it is usually hedging. Also penalise specification by enumeration — parroting back the test cases does not generalise.
-
Generate, run against hidden tests, and iterate on the spec only
Generated code goes straight to the tests. On failure students see the first failing case and may revise the prompt, never the code. Each revision is logged, and the iteration history is a second gradable artifact.
What you only learn by doing it: Require a one-line note with every revision saying what was ambiguous in the previous prompt. Without it students revise by random perturbation until something passes and learn nothing; with it you get a readable trace of requirements reasoning, and it takes thirty seconds.
-
Run the refutation round — hand them passing-looking code that is wrong
Any assistant to generate variants; a code-reading worksheet
Give students generated code for a different problem, some correct and some subtly wrong, and have them predict outputs by hand and justify each verdict. Success predicting outputs was 32.5% overall and only 15% when the code contained a bug, versus 50% for correct code.
What you only learn by doing it: Mix in correct code and do not tell them the ratio. In that study 90.6% of students assumed the code was correct at least once despite evidence otherwise — automation bias is the finding, not a side note. Also brace for unfamiliar idioms: 56.3% reported Python constructs they had never seen.
-
Debrief the deltas: spec ambiguity versus model error versus misreading
Projector and a handful of anonymised submissions
Project three or four real submissions — one precise, one verbose, one specified by enumeration — and have the class predict what the model would do with each. Then reveal.
What you only learn by doing it: Ask permission and anonymise, but do use real student work. Students detect and discount instructor-constructed bad examples instantly.
Where this breaks down
The model will sometimes produce working code from a genuinely ambiguous specification, which rewards a bad prompt and quietly undermines the lesson. Run your own prompts through first and discard any problem where sloppy wording still passes.
This teaches specification and comprehension well and teaches code writing not at all. If your course or articulation agreement requires from-scratch fluency, this supplements rather than replaces it — and that is a departmental decision to make deliberately rather than by drift.
Watch for a language-equity effect. The comprehension study found students raised speaking languages other than English scored significantly lower on prompt comprehension (68% versus 88%). In a community college classroom that is not a footnote — scaffold the reading and grade the specification's content rather than its English.
Provenance: documented practice. The visual-problem exercise is Denny, Leinonen, Prather, Luxton-Reilly et al., “Prompt Problems,” SIGCSE TS 2024. The curricular framing is Vadaparty, Zingaro, Porter et al., “CS1-LLM,” ITiCSE 2024. Every figure in steps 4 and 5 is from “Beginners Struggle to Understand LLM-Generated Code” (n=32).