The semester the AI tools arrived, my problem sets stopped measuring anything. Not gradually. All at once. Any student with a free account could produce correct solutions to every standard intro-physics problem I could assign, with neat explanations besides. My colleagues had the same season in essays, in code, in lab reports.
The tempting reading is that AI broke education. I don't believe that, and I was in the room. What AI broke was a set of proxies. The problem set was never the point. It stood in for practice, a way to make thinking visible so I could grade it. The essay stood in for structured thought. When a machine can produce the stand-in on demand, you discover, uncomfortably, how much of school was the stand-in and how little was the thing itself.
Faced with that, education has three options, and I have watched all three tried in real departments.
We can ban the tools. This fails on logistics alone: enforcement poisons trust, and the students most willing to cheat are the least likely to be caught. It fails worse on purpose. These students will graduate into hospitals, firms, and classrooms that run on these tools. A course that pretends AI does not exist is preparing students for a world that stopped existing.
We can try to detect our way back to 2019. The detectors are unreliable in published testing, and their false accusations land on exactly the wrong students: the anxious, the rule-followers, and disproportionately the students writing in a second language. An arms race against your own students is a strange definition of teaching.
Or we can redesign. Change the work itself so that using AI well is the skill being taught and measured. Move assessment from the artifact to the process: drafts, checkpoints, oral defenses, work that carries the fingerprints of its maker. Move curriculum from producing answers to what answers cost now: framing the problem, decomposing it, judging the output, knowing what to ask next. Those were always the durable skills. AI just made the price of pretending otherwise visible.
Redesign has a second half, and it is the half I care most about.
For forty years we have known that one-on-one adaptive teaching wildly outperforms uniform instruction. Benjamin Bloom measured the gap at two standard deviations and called finding a scalable substitute the "two sigma problem." We never solved it. We couldn't afford to. Adaptation was a luxury: private tutors for the families who could pay, patience for everyone else.
AI makes adaptation cheap for the first time in the history of teaching. Cheap enough for a nursing student relearning algebra at midnight after a shift. Cheap enough for a seven-year-old whose words come out sideways and whose $300 communication app was configured once, wrong, and never touched again. Cheap enough that "the class moves on without you" can stop being a law of nature.
That is the opportunity, and it comes with a warning label. The same economics that make adaptation cheap make manipulation cheap. Education technology has already shown us what it optimizes when nobody is watching: engagement, streaks, time-on-app, a slot machine with a curriculum attached. An adaptive system that knows exactly where a child is weak is also a system that knows exactly how to keep that child hooked. The redesign cannot be left to whoever gets there first with the fewest scruples.
ThinkAthena is my attempt to do it on purpose. It is a small company with four pillars, in a deliberate order.
Tools for kids the system was never built for, free, forever, because that is where the need is sharpest and the market most broken. A suite for bright kids pacing ahead of their classrooms, because coasting is its own quiet damage, and because this is the business that funds the rest. A way back in for college students one frightening course from quitting, because I teach those students and I am tired of watching preventable exits. And eventually, tools for the institutions that teach most of the world's students on the thinnest budgets, because the new economics of software should reach them, and not just the flagships.
Underneath all four runs one engine and one loop: find where the learner actually is, prescribe the next right thing, watch honestly, adjust. And underneath the engine, a floor we will not drill through: never optimize for addiction, never sell a child's data, never claim what the evidence cannot back, never ship a feature that raises engagement while lowering learning.
I am a physics professor. I spend my mornings with students the system is failing in slow motion, and my evenings building tools I can hand them. This site is where I show the work: the essays, the prototypes, the decisions, the mistakes. Learning will be redesigned for the age of AI either way. The question is whether it happens on purpose, by people who teach, or by accident, by people who A/B test.
I vote on purpose. Come watch, or come help.
- Weber-Wulff et al., "Testing of detection tools for AI-generated text", International Journal for Educational Integrity (2023)
- Liang et al., "GPT detectors are biased against non-native English writers", Patterns (2023)
- Bloom, "The 2 Sigma Problem", Educational Researcher (1984)