Assigning an AI build: outcomes, alignment, and an honest evidence base

Intermediate ⏱ 60 min 📅 Oct 2, 2026

Students can now produce a working, shareable piece of software without being able to write it. The instinct is to treat that as either cheating or magic. It is neither. The build is fast; getting it to actually work is where the effort moved. That middle part — use it, notice precisely what is wrong, describe the defect so someone else could reproduce it, repeat — is a real professional skill with a research literature behind it, and unlike the finished artifact, it is assessable.

This is a three-week project you can run in almost any discipline, with the outcomes, the alignment, the rubric, and an honest account of where the evidence is strong and where it is not.

Learning outcomes

Written as outcomes rather than activities — observable, assessable, and at the analysis-and-evaluation level the ASCCC glossary describes. On completing this project, students will be able to:

  1. Translate an open-ended brief into a specification of three to six testable requirements that a third party could verify.
  2. Evaluate an AI-generated design proposal against that specification and identify at least two substantive discrepancies before approving it.
  3. Locate and describe a defect in a form another person can reproduce — steps taken, expected result, actual result.
  4. Verify AI-generated output against an external standard and revise when it fails, rather than accepting output that appears plausible.
  5. Assess their own artifact against WCAG 2.1 AA criteria and remediate what fails.
  6. Compare their predicted performance against the rubric and the actual outcome, and account for the difference.

Alignment

OutcomeActivityEvidence collected
1 — SpecificationWritten requirements, peer-reviewed for testability, before any tool is openedThe spec, with peer review notes
2 — Evaluate a proposalRequire a design plan before any code; mark it up before approvingAnnotated plan showing what they changed
3 — Describe a defectPlay-test, log defects in a fixed format, one per entryDefect log, minimum eight entries with outcomes
4 — VerifyCheck the build point by point against the specVerification pass with results
5 — AccessibilityKeyboard, labels, contrast and colour-independence passAccessibility checklist with remediation notes
6 — CalibratePredict the rubric score and the first user's difficulty; compare afterwardsPrediction, outcome, and written account of the gap

Rubric

Weighted toward the defect log on purpose. In a seven-institution study of novice debuggers, students located only about 70% of the bugs in a program but fixed 97% of the bugs they managed to locate — the bottleneck is finding and describing, not editing. Professional defect reporting research points the same way: when developers were surveyed about what makes a bug report usable, steps to reproduce came top at 83%, and those are precisely the items reporters find hardest to supply.

CriterionWeightDevelopingProficientAdvanced
Specification15%Wish list; untestable language3–6 requirements a third party could verifyRequirements anticipate edge conditions and state what is out of scope
Critical reading of the plan15%Approved essentially unchangedTwo or more substantive discrepancies caught and correctedCatches an assumption that would have surfaced only after the build
Defect log35%Vague reports; cannot be reproduced from the entryEight or more entries with steps, expected, actual, outcomeEntries isolate the condition; failed fixes are recorded as honestly as successes
Verification and accessibility20%Spot checks; accessibility untouchedSpec checked point by point; keyboard, labels, contrast, colour-independence remediatedFinds a failure the checklist does not name and fixes it
Working artifact and calibration15%Works only on the build machine; prediction absent or unexaminedLive and usable by a stranger; prediction compared against outcomeAccounts for the gap in terms that would change their next estimate

Three weeks

Week one — specification. Requirements written in class, on paper, before anyone opens a tool. Peer review for testability. Then the design conversation and the annotated plan, due end of week.

Week two — build and defect loop. Smallest running version, then play-testing. Run a swap day: students test each other's builds and file defect reports against them in the required format. Students find other people's defects far faster than their own, and filing against a stranger's work forces the precision the rubric measures.

Week three — verify, remediate, ship. Verification pass, accessibility pass, publish, cold test with someone uninvolved. Prediction recorded before the cold test, compared after.

Before you assign it: four decisions

Which tool, and who pays. The joint CCCCO, ASCCC and CIO memo of 3 August 2026 is direct: reliance on free consumer tools or personal subscriptions "should not be the basis for required coursework," and faculty "should not be expected to require AI use until appropriate institutional tools are available." Use what your college supports. Requiring a paid subscription also makes the section ineligible for Zero Textbook Cost designation.

An opt-out path. The system's HUMANS framework states that people "should be able to opt out, where appropriate." Decide in advance what the alternative route to the same six outcomes looks like — a specification-and-defect project against existing software reaches most of them.

Age. Dual-enrolment students frequently cannot open these accounts under the terms of service, and you cannot consent on their behalf. Team structures or instructor-hosted tooling solve it; a surprised roster in week one does not.

What goes public, and where. Title 5 §55200(c) requires equivalent access "with substantially equivalent ease of use," and the Chancellor's Office frames that as applying on the first day of class rather than on request. Whether a student-published artifact falls under the federal Title II web rule is unresolved — the rule's third-party exception does not address coursework, and the exposure is higher when the hosting is something the course arranged. Putting WCAG 2.1 AA in the rubric converts that ambiguity into outcome 5.

Why this design, and where the evidence is thinner than people claim

Specification first rests on expert–novice design research: observing practising engineers and students on the same task, the experts spent substantially more effort on problem scoping. That is an expert–novice comparison, not an intervention study — it locates the difference, it does not prove that teaching spec-writing causes better outcomes.

Struggle before consolidation is supported for concepts and transfer: a meta-analysis of 53 studies puts problem-solving-before-instruction at g = 0.36 for conceptual knowledge and transfer. Two honest qualifications. The same analysis finds g = −0.03 for procedural knowledge, so do not make students struggle to discover tool mechanics — teach those directly. And the undergraduate subgroup is smaller, g = 0.28. There is a related finding worth knowing: invention-before-instruction outperformed direct instruction specifically when students had resources available during assessment, which is exactly the condition of working with AI.

Requiring explanation is the best-supported element here. A meta-analysis of 69 effect sizes puts self-explanation prompting at g = 0.55, robust across subject areas and education levels, and a separate study found the transfer benefit came from the explanation requirement independent of how students were taught.

Project-based learning: state the range, not the best number. One meta-analysis of 46 effect sizes reports d = 0.71 for achievement and finds education stage not significant. Another, covering 66 studies, reports an overall SMD of 0.441 but a university subgroup of just 0.116, and says plainly that the effect at university level is relatively low. Those two disagree. I have found no meta-analysis or rigorous causal study of project-based learning at community colleges specifically, and the broader community college CTE literature is characterised as short-term and mixed. Claim that this is a well-constructed authentic project; do not claim a community college evidence base.

Grading process rather than product is a validity argument, not an outcome finding. The case — that what deserves assessing is the student's capacity to judge quality — is a position in the assessment literature, and a good one. I did not find meta-analytic or causal evidence that grading process artifacts improves learning. Say "defensible on validity grounds."

Calibration needs an external standard, or it does nothing. This is where I changed the design. Asking students to predict and then see the result is close to useless on its own: across thirteen exams in a single course, students' overconfidence did not decline at all, and prior scores were unrelated to their later predictions. A meta-analysis of calibration interventions puts the whole literature at g = 0.25, finds that what works targets external standards, and finds that manipulating only the timing of self-judgement actively hurt accuracy. So outcome 6 requires comparison against the rubric and the observed outcome, not prediction alone. And drop the Dunning-Kruger framing — it does not appear in this literature, which has better-grounded constructs.

Why verification is graded separately. In a controlled study, participants writing code with an AI assistant produced less secure code while being more likely to believe it was secure; participants who distrusted the assistant and reworked their prompts produced fewer vulnerabilities. Observational work with CS1 novices found weaker students over-relying on incorrect suggestions, with what the authors termed an illusion of competence, and a study of 19 novices using Copilot documented a false sense of progress from simply having code on the screen. Verification does not happen unless you assess it.

Debugging is teachable and under-taught. A 2024 systematic review of debugging interventions concludes the area remains underexplored and that instructors often lack awareness of effective methods; a controlled classroom study found explicit instruction in a systematic debugging process improved both performance and self-efficacy over practice alone, though with a small sample. The ACM task force on generative AI and programming assessment, reporting in February 2026, points in the same direction: more emphasis on debugging and testing, and requiring students to explain all code including AI-generated portions.

Beyond computing

DisciplineArtifactWhat the defect log captures
MarketingLanding page for a real local client with a working enquiry formForm failures, mobile layout, unclear calls to action
Graphic designInteractive style guide a developer could build fromSpecification gaps found by someone trying to use it
English and ESLSelf-scoring practice activityFeedback text that misfires on correct answers
BusinessBreak-even or pricing calculatorArithmetic edge cases, unhandled inputs
Child developmentClassroom routine or transition tool for a stated age bandAge-appropriateness failures found in observation
Automotive and tradesDiagnostic decision-tree reference for one systemBranches that dead-end or contradict the service information

The brief changes; the six outcomes do not.

References

  • Atman, C. J., et al. (2007). Engineering design processes: a comparison of students and expert practitioners. Journal of Engineering Education, 96(4), 359–379. doi:10.1002/j.2168-9830.2007.tb00945.x
  • Bettenburg, N., et al. (2008). What makes a good bug report? SIGSOFT/FSE-16, 308–318. Expanded as Zimmermann, T., et al. (2010), IEEE TSE, 36(5), 618–643. doi:10.1109/TSE.2010.63
  • Bisra, K., et al. (2018). Inducing self-explanation: a meta-analysis. Educational Psychology Review, 30(3), 703–725. doi:10.1007/s10648-018-9434-x
  • Chen, C.-H., & Yang, Y.-C. (2019). Revisiting the effects of project-based learning on students' academic achievement. Educational Research Review, 26, 71–81. doi:10.1016/j.edurev.2018.11.001
  • Denny, P., et al. (2024). Prompt Problems: a new programming exercise for the generative AI era. SIGCSE 2024, 296–302. doi:10.1145/3626252.3630909
  • Fitzgerald, S., et al. (2008). Debugging: finding, fixing and flailing. Computer Science Education, 18(2), 93–116. doi:10.1080/08993400802114508
  • Foster, N. L., et al. (2017). Even after thirteen class exams, students are still overconfident. Metacognition and Learning, 12(1), 1–19. doi:10.1007/s11409-016-9158-6
  • Gordon, S., et al. (2026). ACM Task Force on Generative AI and Programming Assessment: Final Report. acm-education-genai-task-force.github.io
  • Janssen, N., & Lazonder, A. W. (2024). Meta-analysis of interventions for monitoring accuracy in problem solving. Educational Psychology Review, 36, 96. doi:10.1007/s10648-024-09936-4
  • Michaeli, T., & Romeike, R. (2019). Improving debugging skills in the classroom. WiPSCE '19.
  • Perry, N., et al. (2023). Do users write more insecure code with AI assistants? ACM CCS '23, 2785–2799. arXiv:2211.03622
  • Prather, J., et al. (2023). "It's weird that it knows what I want": usability and interactions with Copilot for novice programmers. ACM TOCHI. doi:10.1145/3617367
  • Prather, J., et al. (2024). The widening gap: the benefits and harms of generative AI for novice programmers. ICER 2024. doi:10.1145/3632620.3671116
  • Schwartz, D. L., & Martin, T. (2004). Inventing to prepare for future learning. Cognition and Instruction, 22(2), 129–184.
  • Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works. Review of Educational Research, 91(5), 761–798. doi:10.3102/00346543211019105
  • Yang, S., et al. (2024). Decoding debugging instruction: a systematic literature review. ACM TOCE, 24(4). doi:10.1145/3690652
  • Zhang, L., & Ma, Y. (2023). A study of the impact of project-based learning on student learning effects. Frontiers in Psychology, 14, 1202728. doi:10.3389/fpsyg.2023.1202728
  • CCCCO/ASCCC/CCCCIO (2026). AI in Teaching and Learning: Building Local Frameworks, Memo ESS 26-52, 3 August 2026.
  • CCCCO (2024). Generative AI and the Future of Teaching and Learning, Report to the Board of Governors, 17 July 2024 (HUMANS framework).
  • CCCCO (2023). Guidance for Distance Education Regulation Changes, Memo ESS 23-14; Title 5 §55200(c).
  • ASCCC (2019). Student Learning Outcomes Glossary.

Every source above was checked against a publisher, repository or issuing-body page before it was cited here. Where two meta-analyses disagree, both are reported. Where I could not find evidence for a design choice, the text says so rather than implying support.