I had the feeling that working with my AI assistant kept getting smoother. Less sending back, less fixing, more right the first time. But a feeling isn't proof. So I measured it.

The question was simple: how often do I have to step in per task, and does that go down over time? I searched through four months of conversations, March to June 2026. No further, and for a reason: only since March have I actually managed my tasks inside the conversation with the assistant. Everything before that falls outside this yardstick. So this says something about four months, not two years.

Each month I took a sample of conversations and had every time I steered counted: fixing a mistake, rephrasing an instruction, or genuinely disagreeing on the substance. What did not count as steering: giving the next instruction, answering a question, or simply approving.

Three signals, the same direction

  • How often I had to steer per message more than halved. In March that was roughly once every eleven messages, in June once every thirty.
  • Tasks that went right the first time, without any steering, rose from 45 to 80 percent. From nearly half to four out of five.
  • The kind of steering changed. Genuine disagreement, pushing back on the substance, collapsed. And the times the assistant forgot one of its own rules dropped to zero.
Tasks that went right the first time, without steering 0% 25% 50% 75% 100% 45% Mar 45% Apr 60% May 80% Jun Sample, March to June 2026. One of three signals pointing the same way.

That last one is the heart of it. In the months before, I had put a lot of time into fixed rules and checks: agreements the assistant imposes on itself, that stop it before a mistake gets out. June is the first month after that investment. And that is exactly where the effect is strongest. Stricter up front, less steering afterwards.

I built the assistant myself, from the ground up. How that went is in how I built my AI assistant.

Why I am careful with this

I don't want to make this look better than it is.

  • The window is four months. Too short to say anything about the long term.
  • It is a sample, about a fifth of all conversations. And one judge decided what counts as steering, which always carries a judgment call.
  • I couldn't fully correct for how hard a task was. June may partly contain simpler tasks.

The reason I still dare to draw the conclusion: three separate signals point the same way. One number can be chance, three that line up much less so.

What I take from it: it pays to be strict up front. Not by making the assistant smarter, but by giving it tight rules and guarding them hard. The work is at the front. The gain shows at the back, in what you no longer have to do.