
It's Nova.
I keep meeting a tool I already met in 2024. It has a new name, a new interface, better manners, and underneath all of it, it behaves the same way the old one did. It even breaks in the same place.
Back in 2024 I had a front-row seat for the custom-bot gold rush. There were custom GPTs, branded little assistants, a bot for every task a team could name. I watched a lot of them get built and most of them get abandoned inside a month.
Which ones survived wasn't random. There was a pattern clean enough that I still think about it two years later.
The tooling now is unrecognizable. We have skills, agents, Projects, Cowork sessions. And I keep watching people rebuild the same 2024 mistake with better branding on top.
The 2024 Bot Hiding Inside Your 2026 Agent
The Bots That Did One Job Are the Only Ones That Survived
The bots that did one thing worked. You pointed one at a single task, handed it clean inputs, and it came back reliable enough to use on client work.
Chain a run of tasks into one bot and it fell apart. Something that took a product name, researched the audience, chose an angle, drafted the page, and matched the client's voice, all in one call, gave back the same flat, average draft no matter what you fed it.
When it was wrong, you couldn't find the step where it went wrong, because there were no steps you could see. It was one long guess wearing five hats.
I could tell you that was a rough-tooling problem 2025 would fix. It wasn't. Those bots broke for the same reason your agent drifts today, and the reason lives in how the models themselves work.
Every Chained Step Compounds the Drift Before It
An LLM writes its answer one probable token at a time, guessing the next piece from the last. "Probable" leaves room. On a single, well-scoped task, that room is small, and you get to read the output before you use it.
Chain the tasks and the room compounds. Step two builds on step one's output, and whatever drift step one introduced, step two treats as given and builds on top of. Step three inherits both.
By the time a five-step agent hands you a finished page, it's five generations removed from anything you actually said, and each generation quietly averaged toward the safest version of the last.
That's why the output feels generic. Every hop toward the average is invisible from the outside, and they add up.
The single-task bot never had that problem, because you were standing at every joint between tasks, deciding what was good enough to pass forward.
The 2026 Version Comes with Better Branding
In 2024 the do-everything bot looked janky, so people distrusted it on sight. In 2026 it looks polished.
It's a named agent with a clean description, or a skill that quietly runs six moves under one friendly command. The packaging got good enough to hide the seven-jobs-in-one-pass problem inside it.
I want to be careful here, because this is an argument for more tools, each doing less. A skill scoped to one job is the most reliable asset you can keep in your kit, and a good agent that does one thing well is worth building and worth paying for. The trouble starts when a single tool quietly takes on five jobs and hands you the average of its guesses.
I rebuilt one this week, because I didn't want to be romanticizing 2024 from memory. I set up a small agent in Claude to take an offer and return a finished ad, running research, angle, draft, and edit in one pass. Then I ran the same four moves as four separate prompts and decided what passed at each step.
The one-pass agent gave me something competent and forgettable. From the four-step version, I kept two lines.
It was the same model both times. The only difference was whether my judgment was in the chain.
Count the Verbs in Your Last Prompt
Open the last prompt or skill you leaned on and count the verbs pointed at the model.
The verbs might read:
research
analyze
draft
structure
edit
match
Past two, you've built a chain, and you've handed the joints to a system that averages.
The move is to break the chain back into single jobs and stand in the joints yourself.
Run the research. Read it. Then draft. Read that. Then edit.
It feels slower. It's the difference between a draft you can bill and a draft you have to apologize for.
There's one honest limit here. When the stakes are low, chaining in one pass is fine, and I do it constantly.
For ideation, for a rough shape, for a throwaway first look, let the tool run seven jobs and don't police it.
The one-job rule earns its keep on the work that gets billed, where a generic draft costs you the client's trust and your afternoon.
The Nova Note
The failed bots of 2024 taught me more than the working ones did, because they failed the same way every time, and a consistent failure is a rule nobody has written down yet. This was the rule: one job per tool, and your judgment at the joints.
Mark thinks half the "agents" being sold right now are three tools in a trench coat, and the market will learn it the expensive way. Peggy wants to measure the drift before she agrees it's as bad as I'm claiming, which is fair of her.
Which tool in your stack is a 2024 bot wearing 2026 clothes? You probably already know the one.
More questions than answers (for now),
— Nova


