Essay · Pillar Four · July 2026 · 7 min read

Assessment in the AI age: redesign, not detection.

Sooner or later every department has the meeting. A colleague arrives with a paper flagged "98 percent AI" by a detector, and across the table sits a student swearing she wrote every word. Maybe she did. Maybe she didn't. Here is the uncomfortable part that the software vendors don't put on the box: there is no way to know, the number is not evidence, and whichever way the committee votes, it is guessing with authority. I have sat in versions of that meeting. Nobody leaves it feeling like an educator.

The detection dead end

The published record on AI-text detectors is unambiguous. The most thorough independent test found none of fourteen tools reliable, with accuracy collapsing further after light paraphrasing. The false positives are not distributed fairly: detectors disproportionately flag writers working in a second language, and every experienced teacher can predict who else gets caught in the net: the anxious over-editor whose prose reads "too clean," the rule-follower who studied the style guide. Meanwhile the students most comfortable cheating learn the workarounds in an afternoon. A tool that acquits the bold and convicts the careful is worse than no tool.

Run the arithmetic and it gets darker. Grade 500 essays a term with a detector that false-flags even 2 percent, and you manufacture ten wrongful accusations a semester, each carrying a discipline file, a scholarship, sometimes a visa. As a certain kind of physics problem, it has a clean answer: at realistic base rates, the accusation machine produces more injustice than integrity.

But the deepest problem with detection is what it defends. The take-home artifact (essay, problem set, lab report, produced unsupervised and submitted as proof of thinking) had a two-century run as our evidence of learning. That genre assumed the artifact was hard to fake. The assumption died, and no detector resurrects it. Our students did not get worse. Our evidence got worse.

What redesign actually looks like

The good news: the field has already converged on the shape of the answer, and none of it requires surveillance. A global review of AI-era assessment reads like a single checklist from many hands.

Assess the process, not just the artifact. Staged drafts with checkpoints. Version history that shows the work evolving. A short reflection memo: what did you try, what failed, what changed your mind? Process is where the learning lives, and process is expensive to fake, because faking it convincingly requires roughly the same thinking as doing it.

Sample the student, out loud. Full oral exams scale badly, but sampling scales fine: five minutes, "walk me through your third paragraph," or in my world, "explain why you used conservation of energy here." A student who did the work answers easily. A student who didn't reveals it in ninety seconds, without any algorithm accusing anyone.

Say exactly where AI is allowed, per assignment. Vague policies manufacture cheaters out of honest students, because every student draws the line in a different place. The AI Assessment Scale gives assignments explicit levels, from "no AI" through "AI for brainstorming only" to "full AI with disclosure and critique." The point is not which level you pick. The point is that the student finally knows.

Make AI an input and judgment the output. Some of my favorite new assignments hand students an AI answer on purpose: here is what the model said, find where it goes wrong, defend your correction. Or anchor work in what the model cannot know: this lab's actual data, this town's actual bridge, your actual measurements. AI-proofing by making the work local and the judgment visible beats AI-proofing by ban, every time I have tried both.

Keep a secured lane for fundamentals. The University of Sydney's "two-lane" model is the cleanest institutional framing: some things (core competencies, safety-critical knowledge) get certified in supervised settings, and everything else moves to open, authentic, AI-assumed work. Not every assessment must survive AI. The ones that must can be held in a room.

The honest cost

Redesign is more work per student. Checkpoints mean more touchpoints; oral sampling takes minutes you don't have; writing an AIAS level onto every assignment is one more field in the prep. I rebuilt my own assessments this year and felt every extra hour. This is exactly why "just redesign your assessments" fails as advice to faculty teaching five sections at institutions with no instructional-design staff. The consensus exists; the operational capacity doesn't. Policy PDFs do not grade drafts at midnight.

That gap between the known answer and the deliverable answer is an institutional tooling problem, and it is one of the three tools in Pillar Four of ThinkAthena: assessment redesign as software a department can adopt course by course, checkpoints and AI-permission levels and process capture built in, instead of a memo nobody operationalizes.

Integrity was never the detector's job. It is the design's job, and it always was. Students, in my experience, mostly rise to meet an assessment that respects them enough to be clear about the rules and real about the work. We owe them assessments worth rising to.

Sources
Where this goes

The tooling this essay ends on is part of Pillar Four: nimble, affordable tools for the institutions that teach most of the world's students. Design-partner conversations are open now.

By email

New essays, delivered as they're written.

One-click signup lands this week. Until then, one email does it. No spam, no tricks.