[Headless AI]

It has either been a year, or it hasn’t, since the output of AI tools crossed a certain quality threshold and convinced even experienced developers that something wicked this way comes.

For a while now we have been working on reshaping Sobamail’s software development processes around AI. Observing the rate of change in some of these processes slow down quite a bit, I decided to write this post to share what we have learned and to create some common ground for exchanging ideas.

First, this much is clear: Software is still software. If you neglect what the job demands because of time pressure, you are still getting “future you” into trouble. You still need to clean up after the code you produce, regularly. The fight against complexity goes on with the same determination.

So yes, large language model technology, when used right, seriously speeds up how fast your team gets to working code. But it does not put you on the right path. Nobody should expect that from a machine programmed for sycophancy anyway.

I could say we approached our new toy with some caution, but when I talk to others, the story I hear is more or less the same. Producing patches by copy-pasting code from the web interface, and later by having IDE integrations write small functions here and there, has today given way to a world where end-to-end features are implemented entirely by the language model.

So what does it take for the model not to lose its way during long-running work? Or, more technically, how should the contents of the context given to the language model be managed?

At first we thought the answer was to split the work into pieces that fit into a single context, read everything the model wrote and keep checking its every step, like the small corrections you make to stay in your lane while driving on a straight road. But on the first big job it became very clear that this torment was not bearable. If the whole working day of an engineer who has spent years getting better at their craft was going to go into holding the language model’s hand, there could be no talk of using resources efficiently.

If a software project is a tree, the code is its leaves.

Building on this principle, I got into a meta-discussion with Opus, and what came out of it was a skill the model calls “handoff looping”.

To start a handoff loop, you hand the work to the model as a paragraph of text. Guided by that paragraph, the model takes a tour of your existing code (it calls this “measurement”), examines your paragraph in light of what your code offers, and identifies what is missing. It then comes back to you with a few .md documents (it calls these “loop documents”) that explain in detail how it will close those gaps piece by piece (it calls the pieces “passes”), and a shell script (it calls this the “driver”).

Next comes the meeting phase (what the model calls the “signoff round”), where the model asks you its questions. As you answer them (it calls the answers “owner verdicts”), it writes them into the loop documents, and a work plan emerges.

Once your meeting with the model is over, you run the driver script: the engineer’s shift ends and the model’s begins. The driver runs the model again and again with the same prompt, making it reread the documents it keeps updating as it makes progress. After finishing each piece, the model commits its work, updates its documents to point at the next piece of work, and ends that session. The result is a setup where the model always works on the next job with a clean context, until it either gets stuck somewhere and gives up, or finishes the work and stops.

It is possible to have models run like this for days. Every now and then you can also check on the work by telling another model “read the docs and give me a summary of where our loop stands”. I never recommend reading the model’s documents yourself. Having a fresh model summarize the current state is a far more efficient use of your time.

I will update this post when I manage to put the skill up on Github.