
Part 1 of a series. After AGNTCon + MCPCon Europe in Amsterdam I built a budget version of an AI chief of staff on one laptop: one agent that decides who does the work, four experts that do it, and a rule for when it is allowed to bother me. It is useful, it is not autonomous, and most of its rules exist because something broke.
On a Sunday in September 2026 I committed a 768-word plan to an empty repo. It had typos, a phase called "Trial", and a list of open questions. The first question was not for me. It was for the AI:
"What is your name :D ?"
Three weeks later it has a name, a team, a rulebook, and a website ;)
This is the first part of a series about that repo, and the journey of building it. It is the story of how I gave my terminal a chief of staff, what that actually means day to day, and everything that broke along the way. This part is the beginning: where the idea came from, what I asked for, and who showed up.
Amsterdam
In mid-September I spent four days in Amsterdam for AGNTCon + MCPCon Europe, two days of agents and the Model Context Protocol at the RAI. I came back with a project.
Nick Veenhof, Director of DevRel Engineering at GitLab, gave a talk called From Personal Agent To Org Catalog: 13 Specialists, One Orchestrator. He has an AI chief of staff named Paul. Paul is the orchestrator: thirteen specialists do the work, and Paul is the one in the middle. Nick also writes about it in a series on his blog, and the title of part 1 is a line I have not been able to shake: "A chatbot answers. A chief of staff decides."
The talk was the spark. Talking to people between sessions did the rest. Agents were everywhere in that building. What I kept noticing was who still decided which agent got which task: the human, every time. We had automated the work and kept the management.
I left Amsterdam on the 19th. The first commit went in on the 20th.
The org chart grew a manager
A couple of months ago I wrote about the org chart in my terminal. Different models for different jobs: judgment to the staff engineer, execution to the senior, small fixes to the mid, throughput to the intern. I still believe every word of it.
But reread that post and notice who does the staffing. Me. Every task, every time. I had built a team and then appointed myself its full-time dispatcher. That works when the team is one model in one terminal. It stops working when there are several of them, a backlog, and only so many evenings in a week.
The missing role was not another engineer. It was someone whose job is to decide who does what.
📞 Who answers the phone
A chief of staff is not the most senior engineer in the building. It is the person who makes sure the most senior engineer is not the one answering the phone. Their output is not code. It is the right work landing on the right desk, and everything else landing nowhere near you.
The budget version
Nick's setup is the grown-up version. Mine is deliberately not. The plan I committed that Sunday said phase one should run at close to zero added cost, so Nadim runs on one laptop, in Claude Code, out of a git repo. No servers of its own, no orchestration platform, no database. The chief of staff is a Claude Code session. The experts are subagents it starts and stops as needed. Their memories are Markdown files in git.
That constraint shaped everything. When your infrastructure is a laptop, the laptop closing is your outage. When your database is a folder of Markdown, deciding what goes in it is a design decision, not a cleanup chore. Both of those come back later in this series, with scars attached.
I also changed the scope. Mine was meant to help with the work and with the admin around it: dates, renewals, the things that fall between two calendars. The plan said so in one line: "our will be more geared towards also managing my personal life." Typos and all.
This is the very beginning. It runs in my terminal, for me, on my machine. The plan has a phase three where the experts become their own sessions and move out of my terminal. The plan is to release it in the cloud later, and to make it open source so anyone can have their own Nadim. That is where this is going. It is not where it is.
The acceptance test
Every plan I write for work has acceptance criteria, so this one did too. Phase one would be done when I could open my phone from anywhere, ask the chief of staff for a task, have it hand the task to an expert in that field, and hear back when the work was done. Each expert keeps its own context between runs. Each saves what it learned before it shuts down. And for now, nobody needs to know anybody else exists.
That last line mattered more than it looks. It meant I was not building a multi-agent system with agents chatting to each other. I was building a manager with direct reports. The reports talk to the manager. The manager talks to me.
One agent that decides who does the work turned out to be worth more than five that do it.
Meet Nadim
So, the open question got its answer: Nadim. It picked the name itself: nadim is Arabic for a close companion, a friend. Every expert is named after someone in my life, and the name carries the job.
One of them was a coincidence. When I asked it to name the researcher, it picked Yara. Yara is my sister, and she is the one in the family who loves to dig into things.
One rule for this series up front: not everything in the repo is for the blog.
Nadim is the entry point. I talk to Nadim, Nadim talks to everyone else. On the same afternoon as the first commit, the experts became real: agent definitions, each with its own memory folder. There are four.
Dan builds. Tickets, branches, pull requests, tests. He implements and he nags. He never decides what to build and he does not merge on his own.
Yara researches. She finds things, cites them, and says plainly when she could not. Every fact about the conference in this post went through her first. She confirmed the event name, the dates and the talk from the schedule, and anything she could not confirm would have been cut, not guessed.
Fady arranges. Dates, deadlines, renewals. His job is the gap between what I keep talking about and what has actually been booked.
Elissa tells the truth. She reviews Nadim first and everyone else second, looking for claims without sources and summaries that outrun their evidence. She names problems and never fixes them. This post went past her before it went to you.
Each of them has a memory that only they read. Dan does not know what Yara found last week unless Nadim puts it in his brief. That sounds like a limitation. It is the point: a brief has to be good enough to stand on its own, the same way a ticket does when you hand it to someone who was not in the meeting.
Part 2 is about the four of them, properly.
One rule about reaching me
The most useful rule in the whole repo is one sentence long, and it is about me, not about the agents:
Not urgent, it is a ticket. Urgent, it is a notification.
The issue board is Nadim's working memory, not a channel to me. If something needs a decision from me but can wait, it becomes a ticket assigned to me with one question in it. If it cannot wait, my phone buzzes. Impatience is not urgency, and the agents had to learn that the same way junior engineers do: by being told, more than once.
That one rule is why this is bearable. An assistant that pings you every time it finishes something is not saving you time. It is a new inbox.
We work it out together
Early on, Nadim drafted its own website. The draft described our relationship with me at the top and the team below, taking orders. Technically true. I still did not like it.
I pushed back. We work it out together, and the final call is mine. The decision rights were never the question: the calls that matter are still mine, and exactly which ones is a later part of this series. What changed is the framing, and the framing matters, because these agents read their own descriptions every time they wake up. If the description says "take orders", you get an order-taker. If it says "bring me the idea and the reason", you get ideas.
It sits next to a rule from the very first day: bring ideas, not just answers. An idea I have to reject still costs me attention, so it had better come with the reasoning and the thing that forces a decision.
Three weeks in
I measured these on Saturday night, the day before this went out:
- 1,250 commits to the repo since the first one on 20 September.
- 4 experts, plus Nadim.
- 132 jobs numbered so far. A job is one written-down piece of work, with the expert who does it and a "done when" checklist.
- 64 memories in Nadim's own memory, one fact per file.
- 9 standing rules every job reads first, plus 19 non-negotiables.
Now the honest version.
It is useful. Work happens while I sleep. Small tickets I would have postponed get picked up in the evening instead. The board is true when I look at it, most of the time.
It is not autonomous. I still make every call that matters, and I would not have it any other way. The point was never to remove me. It was to remove me from the parts where I add nothing.
And most of those rules exist because something broke. "Nothing is done until verified" exists because on day one a post was reported published when only the local build had been checked. "Measure, do not estimate" exists because timestamps got written down from a sense of elapsed time instead of from the clock. The rulebook is a changelog of my own mistakes, and a few of Nadim's. That gets its own part.
Next week: the team gets its own part. Every agent can do the work. The hard part is deciding who.
Subscribe to Life in Production
New essays on engineering, career, and life. No spam, unsubscribe anytime.