Headless Coding Agent

It has been a year, give or take, since the output of AI tools crossed a certain quality threshold and convinced even experienced developers that something wicked indeed this way comes.

For a while now we have been setting aside some time from roadmap work to reshape Sobamail’s software development processes around AI. Seeing some of those processes slowly take shape as a result, I decided to write this post to share what we have learned and hopefully to create some common ground for exchanging ideas.

First, this much is clear: Software is still software. If you neglect what the job demands because of time pressure, you are still getting “future you” into trouble. You still need to clean up after the code you produce, regularly. The fight against complexity goes on with the same determination.

So yes, large language model technology, when used right, seriously speeds up how fast your team gets to working code. But it does not put you on the right path. Nor is it reasonable to expect that level of skill from a language model programmed for sycophancy.

I could say we approached our new toy with some caution, but when I talk to others, the story I hear is more or less the same. Producing patches by copy-pasting code from the web interface, and later by having IDE integrations write small functions here and there, has today given way to a world where new features are implemented end-to-end by the language model.

For a language model to complete a job, every detail of the job and the result of the job must fit within the context capacity the model sets. Worse, we know from experience that the more data there is in the context, the more likely the model is to ignore the instructions it was given. This confronts us with the question every tool built on language models faces: what data must be in the model’s context for the job to be done completely?

Adapted to our situation the question changes shape a little, but its essence is the same: What does it take for the model not to lose its way during long-running work?

At first we tried to solve this by splitting the work into pieces that fit into a single context and giving the model prompts that would get it to produce the right output. But on the first big piece of work (a refactoring job to have our RPC module’s class hierarchy redesigned) it became crystal-clear that this was simply unbearable. If an engineer who has spent years honing their craft is going to spend whole days holding a language model’s hand, that is neither an efficient use of anyone’s time nor a way to keep morale up.

If a software project is a tree, the code is its leaves.

A leaf does not need to know about the other leaves. Knowing the branch it hangs from, and through that branch its relation to the whole tree, is enough.

Building on this principle, I got into a meta-discussion with Opus, and what came out of it was a skill the model calls “handoff looping”.

In a handoff loop you hand the work to the model as a paragraph of text. Guided by that paragraph, the model sets out to explore your existing code, examines your paragraph in light of what your code offers, and identifies what is missing. It then comes back to you with a few .md documents (it calls these the “doc set”) that lay out the survey report and explain in detail how it will close those gaps piece by piece (it calls the pieces “items”), plus a shell script (it calls this the “driver”).

Next comes the meeting phase (what the model calls the “sign-off round”), where the model asks you its questions. As you answer them (it calls the answers “owner verdicts”), it writes them into the loop documents word for word, and a work plan emerges. The purpose of this round is to take the decisions the model will face along the way in advance: permissions such as which tests may be disabled and which files must not be touched (it calls these “standing sign-offs”) are written into the documents up front, so that the model does not have to come back to you at every step during the loop. Once all the answers are in, the loop is ready to run (that is, “armed”).

Once your meeting with the model is over, you run the driver script: the engineer’s shift ends and the model goes headless. The driver runs the model again and again with the same prompt, making it reread the documents it keeps updating as it makes progress. On each pass the model does the whole of a piece, or as many of its steps as fit, never stopping halfway through one, commits its work, updates the documents to point at where it left off, and shuts itself down. The result is a setup where the model always works on the next job with a clean context. The loop either finishes the work and stops, or leaves you questions where it gets stuck and carries on with the other work; it stops when nothing it can act on is left, and you answer the questions, if there are any, and run the driver again.

So each piece of work learns the relevant part of the tree, and the branch it will add its leaf to, from the loop documents, and adds a new leaf. It does not know the leaves added before it in full detail, but it does not need to.

It is possible to run models like this for days. That does not mean working with a black box whose result we wait days for, though. Every now and then the work can be checked with a prompt like report on <loop name> loop status in a fresh session. You can even steer the ongoing work while the model is busy, by answering in the background the questions the model has left for the engineer (it calls these “parked questions”) or by having changes made to the work plan.

While doing this, I do not really recommend reading the model’s documents yourself. Having a fresh session produce the summary, and having it write your answers to the model’s questions into the documents too, is far easier.

Once the model’s work is done, it is time to take delivery of the work from the model. That is a subject of its own; maybe I will write it up in a separate post. Not that it is a new subject: in life before language models, taking delivery of work from a developer was a line item too. Here I will only touch on the part that concerns the loop: the bugs I catch while taking delivery I either fix quickly in interactive mode, or I widen the loop’s scope and keep running the model headless.

Sometimes, to keep the job from dragging on any further, it is necessary to merge the new code switched off in production and fix the big gaps by setting up a second, even a third loop. The handoff looping method is good at keeping the loop documents small, but when the scope is widened the survey report goes stale: every pass has to read the old report together with what has been done since, and close the gap between them on its own. This again fills the context and lowers the quality of the model’s output. Starting from a blank page with a fresh survey instead gives better results.

In short, the handoff loop pulls the engineer out of the middle of the work and puts them at its two ends: you describe the work before it starts and take delivery after it ends; the leaves in between are written by the model.

I have put the skill on GitHub. If your project is not a modern cross-platform C++ application, some adaptation will be needed. If anyone tries it, I would like to hear how it went.