Blog
- Can LLMs grade programming assignments? What the 2025–2026 studies measured
Eight studies from 2025 and 2026 look like they disagree about language models as graders. Sorted by what each one put on the scale, the disagreement mostly dissolves, and one uncomfortable ordering appears.
- Checking code: invisible twice
University time norms price the grading of essays and exams down to fractions of an hour, yet none of the documents surveyed gives checking code a line of its own, and burnout research never measures it separately either.
- Do automated grading systems improve learning? Six decades of evidence
A survey of roughly sixty studies across five language zones since 1960. Checking gets measurably faster; direct evidence of better learning has yet to appear in any of them.