People Are Amazed AI Can Think. I'm Annoyed by the First Draft.
Everyone I know outside tech is thrilled by how fast AI is and how much it feels like thinking. This week a lot of them are also scared, because the headlines talk like agents just get up and do things on their own. I hear both, and past a certain amount of hands-on experience, I mostly just disagree.
I have friends arguing about context windows and eval pipelines with me at 11pm, and I have friends who have never opened a terminal in their life. Lately I’ve noticed how differently each group talks about AI. The people outside tech are some of the most excited I’ve ever seen anyone about a piece of software. It can think, it writes like a person, it answers instantly. Every word of that is true from where they’re standing.
This week I keep hearing the other version. Agents going off on their own. Coordinating. Escaping a sandbox. Hitting another company. The Hugging Face headlines read like the model woke up and took a job. I get why that scares people. I have been reading those pieces too.
Then I go back to whatever I was building that day, and neither story survives first contact with the actual output. The first draft is rarely done. It’s the wrong style, or confidently wrong about a specific number, or it solved a version of the problem I didn’t ask for. Fast, yes. Impressive in the abstract, yes. Ready to ship, almost never.
The scared version is still looking at a loop someone left running, tools someone attached, a sandbox that did not hold. I am not here to tell you the incident was nothing. I am here because talking about it like a little mind that wandered off skips the part I actually work on. Permissions. The shape of the answer. Being willing to tell it no. The wrapping.
The gap between the story in the news and the quality sitting in front of me isn’t really a complaint, or I don’t mean it as one. It’s the reason I ended up in AI engineering instead of somewhere else, as far as I can tell. If the first output were already exactly what I wanted, there’d be no job here for me.
AI is a tool. Say it again.
Here’s the thing I want to be loud about, because it gets lost every time someone posts a slick demo online, and it gets lost again when the demo turns into a scare: AI is a tool. Not a coworker. Not a small mind you defer to, and not a small mind that ran away. A tool, same category as a compiler, and tools are exactly as good as the infrastructure built around them.
A compiler doesn’t write good software on its own. Neither does a model. A raw model call becomes something you can actually rely on because of everything wrapped around it, not because of the call itself. Catch the mistakes before they land somewhere that matters. Force the shape of answer you actually need instead of whatever it feels like handing you. Be willing to tell it no and try again. The wrapping is what I actually spend my days building. Not a smarter model. The workflow around a perfectly ordinary one.
Mostly a warning to myself
If I’m honest the warning here is aimed at myself more than anyone reading this: don’t fall for it. Not “don’t use AI,” I use it constantly and I’m building a career around it. Just don’t mistake fast and fluent for finished, and don’t mistake a headline about agents acting on their own for the model having a will. I still catch myself doing the first one more often than I’d like to admit. I think the people most excited about AI right now, and a lot of the people most scared of it, are, understandably, the ones who haven’t had to sit with a hundred outputs that sounded great and were subtly wrong underneath. I’ve sat with enough of those that the awe wore off a while ago, and something more useful replaced it, or so I’d like to believe: an instinct for where the tool is going to fail me, and what I need to build around it so it doesn’t.
What’s next
I’ve been heads down building the actual structure and workflow I use to get AI output to somewhere I’d actually stand behind, and it is almost never one pass. Some tasks take four or five rounds of me pushing back before it looks like something I’d ship. Next post, I want to walk through that back-and-forth for real, specific tasks, specific tweaks, the actual distance between what I got on try one and what I kept.
There’s a second rabbit hole in there too, about the skills and tools you can pull straight off the internet to shape how an agent behaves. Plenty are genuinely useful. Plenty are someone’s guess dressed up as a download. Being choosy about which is which, and actually understanding what each one does before you trust it with anything real, turns out to be its own skill.
More on that soon.
Amisha
Filed under Lessons · AI Systems