All writing
LessonsSep 14, 20266 min read

The Mystery of Eleven Discussions, and the Eval I Didn't Mean to Run

Eleven classmates, one discussion prompt with no answer key, and almost every reply hit the same three beats. The question was supposed to show variance. The model answered like there was a right one.

I’m taking an algorithms class this term purely because I like algorithms. Nobody’s making me. So when the first discussion board went up asking why understanding algorithms matters and how Python makes them easier to learn, I sat down and actually answered it. For me personally, I see it this way: algorithms give you a structured way to break a big problem into smaller steps, and most of them lean on data structures to track the state of the problem as you go, so learning one sharpens the other. On the Python side, it’s mostly the boilerplate. Less of it than something like Java means you can test an idea fast instead of fighting the type system first and thinking about the logic second.

Then I opened the rest of the board. Eleven other answers to the same open-ended prompt, the kind of question that should get you eleven different takes from eleven people who each learned to code differently.

Instead I got the same three moves in almost every post. bam. bam. BAM. Algorithms are the backbone of computer science, Python’s syntax is clean so you can focus on the logic instead of the syntax, a GeeksforGeeks citation, a line about semicolons and curly braces, a closer about how this all builds problem-solving skills, same order, every time. It was obvious. Not “hm, some of these feel a little similar” obvious, more like every single answer landed in the same three beats in the same order, and I could tell within a sentence or two which ones had gone through ChatGPT first.

The thing that actually got me wasn’t that people used it. I’d have guessed some of them did the second I saw the pattern anyway. The question didn’t even have a right answer. It was a discussion prompt. Why algorithms matter, how Python feels to learn. A board like that is supposed to show variance across different people, different histories, different days even. Instead the answers came back the same for eleven classmates who are not the same person, people from all over who learned to code in completely different places. Same three beats, same order, same citations. It’s weird. The board still read like the same function called eleven times. Same seed, same output, small cosmetic noise in the variable names. It makes me wonder how sensitive the model actually is to the user sitting in front of it, versus how much it’s just returning the nearest “correct” paragraph it has for a prompt that was never supposed to have one.

I spend my actual working hours thinking about exactly this failure mode, just in a different context. When you’re building eval pipelines for a model, one of the first things you check for is whether it collapses toward the same answer no matter how you phrase the question. A model that gives you real variance across runs is telling you something. A model that gives you the same five sentences back every time, dressed up slightly differently, is telling you something else: it converged, and it’s not actually reasoning through the prompt anymore, it’s pattern matching to the nearest thing it’s seen before. A discussion board full of people doing that is the human version of the same collapse.

The people you’re building for are emotional, and impulsive about it. One day you want one thing, the next day you want the opposite, more or less depending on the vibe. Sam Altman has talked about that in interviews: the same person wants a different personality on different days, and almost nobody wants to go set sliders to lock it in. The model has to be sensitive to that. If it answers a discussion question the exact same way for eleven different humans, I’m not sure how it does the harder version, noticing that this specific person, today, needed something else.

Not because anyone’s dumb, most of these answers were perfectly correct. Algorithms are important, Python is readable, both true. But correct and thought-through aren’t the same thing, and you can tell the difference by whether the answer could only have come from this specific person, or whether it could have come from anyone who typed the same prompt into the same tool. The posts that didn’t sound like the others were the ones anchored to something real. Someone tying algorithm efficiency back to a specific bug they fixed in a mobile game they’d actually shipped. Someone admitting they hadn’t used every language so they weren’t totally sure Python was easiest, just their best guess. Those are the two answers I’d remember a week from now, because a real, specific memory doesn’t have a template to fall into.

Last post I said the gap between a fast draft and a finished one didn’t close as the tool got better, it just got quieter. This is that exact thing, caught live, in the lowest-stakes place it could possibly show up. A discussion board nobody outside eleven classmates will ever read. If it’s already this automatic somewhere nobody’s watching, it’s the same reflex sitting under a tweet, a comment, an argument you’re having with a friend about something you actually care about.

Here’s the line I keep coming back to, and it’s thinner than people think. Using AI means you still show up with a take, and the tool sharpens it, checks it, argues with it a little. Letting AI handle something means the take was never yours to begin with, you just signed your name under whatever it produced. A discussion board question that genuinely asks for an opinion only works if someone actually sits with it long enough to have one. Outsource that enough times and you stop noticing you’ve stopped having one.

A few threads I pulled on while writing this and then cut, because each one is really its own post. Whether the way computer science gets taught, heavy on definitions and citations, light on sitting with something ambiguous, is what’s training people to reach for a tool the second a question wants judgment instead of an answer key. Whether “high agency” even means anything once the tool in front of you is built to make the call for you. And whether real judgment can get built any way other than actually doing the thing and living through it, versus reading what doing it would probably be like. More on those soon.

Amisha

Filed under Lessons · AI Systems

Coming
Your Agent Has Amnesia. Here's the Dict That Fakes Memory.
Week 2