
Stop using one model for everything. Match the model to the task the way you staff a team: judgment to the staff engineer, execution to the senior, small fixes to the mid, throughput to the intern. When it goes wrong, it is usually a staffing mistake, not a model failure.
My son is nine months old and he has exactly one job: grow. He is excellent at it. Nobody asks him to file the taxes.
A good team runs on the same principle. The intern sits in the architecture meeting to learn, not to decide. You ask them to document, to summarize, to sanitize the ticket queue. And you do not ask the staff engineer to sanitize tickets. Not because they can't. Because every hour they spend on it is an hour of judgment nobody else on the team can replace.
Somewhere in the last year, my terminal grew an org chart.
The team
I wrote about promoting my AI from code monkey to teammate. The promotion was real. What I did not say is that there is more than one teammate in there, and they are not interchangeable.
Fable is the staff engineer. The big UX changes, the plans that touch everything, the editorial calls on this blog. Anything where the expensive part is judgment, not typing. I do not bring Fable a typo.
Opus is the senior engineer. Once the what is settled, Opus ships the how. Hand over the plan and the implementation comes back solid, without hand-holding.
Sonnet is the mid-level. Small UX changes, contained fixes, the daily work that needs competence but not a summit meeting.
Haiku is the intern. Summaries, boilerplate, the grunt work. Fast, cheap, endlessly eager, and occasionally caught trying to commit straight to a protected branch. God bless him. He is still doing his best.
And Gemini is the designer on contract. Not even from the same company, which is how real teams work too. Every cover on this blog is Gemini's, including the one above this post. I leave the watermark on. A designer signs their work.
🪜 The seniority ladder
A team is a pricing structure. You pay the intern for throughput, the mid for competence, the senior for execution, and the staff engineer for judgment. Salary scales with the cost of their mistakes, not the volume of their output. Model pricing works exactly the same way. You are not paying per token. You are paying for how expensive it is when the answer is wrong.
That reframing is the whole trick. A salary and a token price are the same instrument: insurance against expensive mistakes. Once you see it that way, choosing a model stops being a benchmarks question and becomes a risk question. Not which model scores highest. Which mistake can you afford.
One model for everything fails twice
Most people pick one model and use it for everything. It fails in one of two directions.
The first is sending everything to the big model. It works. It is also the staff engineer sanitizing tickets. You burn the budget and the rate limits on work the intern would do just as well, and when the moment comes that actually needs the judgment, you have spent it on boilerplate.
The second is sending everything to the small model and then blaming it. You asked the intern to redesign the system. The intern gave it a shot, because interns always do. The result was bad, and now the model "isn't good enough." The intern did not fail you. The staffing did.
Read the task before you assign it
The skill is not only prompt engineering. A good brief still matters, but it is the second step. The first is looking at a task before you assign it and asking what it actually needs. Is this judgment? Is this execution? Is this throughput? That is not an AI question. That is the first question every engineering manager learns to ask, usually after getting it wrong a few times with people who deserved better.
I learned it the usual way. I handed Haiku a branch cleanup, judgment work dressed up as grunt work, and he deleted half my branches with the enthusiasm of an intern finally trusted with something real. That one was not his failure either. I staffed judgment work to a throughput hire. The skill is not easy. It just looks easy.
The tooling quietly made managers of all of us. Nobody sent the invite to the meeting, but here we are, running staffing decisions a dozen times a day between coffee refills.
Match the task to the model and everyone on the roster looks brilliant. Match it wrong and you end up firing the intern for your own staffing mistake. The models are fine. The org chart is on you.
Subscribe to Life in Production
New essays on engineering, career, and life. No spam, unsubscribe anytime.