The Model Doesn't Do Anything. Your While Loop Does.
In June 2026 I built two notebook agents and noticed the model barely does anything. It returns JSON. The loop, the memory dict, and the tool schemas did the real work in those systems.
In June 2026 I built two AI agents in a pair of notebooks, and somewhere in the middle of debugging the second one, I had a moment that changed how I think about this work.
The first agent is a CRM Lead Qualifier. Feed it a lead, and it enriches the company info, pulls whatever history exists in the CRM, and spits out a score. The second is an IT Support agent. Point it at a server, and it investigates what’s wrong, restarts a service if it needs to, or escalates to an actual human if the situation calls for it.
On paper these are completely different problems. One is sales ops, the other is infrastructure. Once I had both running, I noticed they were built out of the exact same parts. Same loop. Same pattern. Same shape.
I also noticed something less comfortable: in those two notebooks, the model itself barely does anything.
Here is what I mean. When my CRM agent “looks up a company,” the model does not go look up the company. It can’t. It has no network connection, no API key, no hands to do anything with. The model pauses mid-response and hands back a small piece of structured text, something like this:
{"function": "lookup_domain_info", "arguments": {"domain": "acmecorp.com"}}
The entire contribution is a function name and some arguments, wrapped in JSON. My code is the thing that reads that request, makes the real HTTP call, gets a real result back, and feeds it to the model so it can decide what to do next.
Same story on the IT Support side. When people say “the agent restarted the server,” what actually happened is the model emitted a request to call a function named restart_service. The restart only happened because I had already written a Python function that does the restarting, and a dispatch table connecting the model’s text output to that real function:
AVAILABLE_FUNCTIONS = {
"get_server_health": get_server_health,
"fetch_recent_logs": fetch_recent_logs,
"restart_service": restart_service,
"escalate_to_engineer": escalate_to_engineer,
}
Strip that dictionary away and the model can talk about restarting the server all day long. Nothing happens.
The model proposes. The harness disposes.
Once I saw this, I went back through both notebooks and realized almost everything I’d been mentally filing under “the AI is doing this” was actually something I wrote.
The loop. A single call to the model is stateless. It reads a prompt, generates an answer, and forgets that the conversation ever existed. Any sense of the agent “working through” a multi-step problem is just my code calling the model again and again in a while loop, each time stuffing the growing conversation history back into the prompt so it looks like the model remembers what happened three steps ago. It doesn’t. I’m just showing it the transcript every single time.
The memory. This one sounds the most impressive from the outside and turns out to be the least mysterious up close. In both agents, “memory” is nothing more than a dictionary I update as the agent runs and re-inject into the next prompt. There’s no persistence, no understanding, no internal state carrying forward on its own. It’s a variable.
The tools. The model only knows a function like restart_service exists because I wrote a schema describing its name, its purpose, and what arguments it takes. Write that description vaguely and the model will reach for the wrong tool at the wrong moment, confidently. The tool selection isn’t intelligence, it’s documentation quality.
Put those three things together, the loop, the memory, and the tool schemas, and you basically have the entire architecture of those two agents. None of it is the model being clever. All of it is code I wrote around a model that, left alone, would just answer one question and stop. I am not saying this is the only way to build agents. It is the architecture that showed up as soon as I stopped treating the API call as the product.
It also reframes what I am actually building when people ask me what I do. I am not fine-tuning a smarter brain. I am building the scaffolding that lets a fairly dumb, stateless text generator look like it is taking real action in the world. The model is the commodity here. Anyone can call the same API I am calling. The harness, the loop, the memory design, the tool schemas, that is the actual engineering, and that is the part nobody can just copy from a model card.
I have started catching myself mid-sentence. “The agent restarted the server.” No, it didn’t. The model returned some JSON. My while loop restarted the server. The same split showed up in production on a clinical write path that needed a composite key, and it is why I treat evals like a test suite instead of reading the output and hoping.
I glossed over a problem those notebooks still have: the memory dictionary works fine when a conversation is short, and it breaks down fast once a conversation gets long enough to blow past the context window. At some point you have to decide what the agent is allowed to forget. The genuinely hard part starts there, and that is the post I am writing next.
FAQ
- Do these agents actually execute code or commands?
- No. The model returns structured text, usually JSON, describing the function it wants called. My code reads that request, runs the real function, and feeds the result back. Strip away the dispatch table and nothing happens.
- How am I using the word harness here?
- The code around the model: the while loop that calls it repeatedly, the dispatch table that maps its requests to real functions, the state I re-inject each turn, and the tool schemas that tell it what it is allowed to attempt. In these notebooks, the harness is where the engineering lives.
- Why do the agents seem to remember things?
- They don’t. Models are stateless between calls. “Memory” is state my code persists and re-injects into the next prompt. Every “the agent remembers” is really “someone is replaying state outside the model.”
- How is an agent different from a model in this setup?
- A model answers one prompt and stops. These agents are a loop wrapped around a model: call, run the requested tool, append the result, repeat until there is nothing left to call. Take the loop away and you have a vending machine.
Amisha
Filed under Agents · AI Systems
change my mind
Send me the counterexample. email me