I think we’re choosing AI models wrong
Stop asking which AI is best. Ask what job you would give it.
Most people still ask which AI is best.
I have stopped asking that. The question I keep coming back to instead: if these models worked inside my company, what job would I actually give each one?
Once I started thinking that way, something became hard to unsee. The labs already answered the question. They just did not use the word "org chart."
The labs are publishing staffing policies now
Here is a line from Anthropic’s developer documentation, published this month. Start with Claude Opus 5 for most workloads. Move to Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evaluations on Opus 5 at higher effort still fall short.
Read that as a manager rather than as a developer. That is not a product page. That is a staffing policy. It names a default hire, it names a condition for escalation, and it names a more expensive person you call only after the default has already tried and come up short.
OpenAI does the same thing from the other direction. Open the ChatGPT model picker and you will not see model names at all. You will see Instant, Medium, High, Extra High, Pro. Those are effort settings dressed as product options. What sits underneath depends on your plan: free and Go accounts are talking to GPT-5.6 Luna, the cheapest model in the family. Paid accounts are talking to GPT-5.6 Sol.
Which means a lot of the public opinion about "ChatGPT" was formed on the intake clerk.
The roster, as I would actually staff it
Claude Fable 5.1 is the outside expert you escalate to. Anthropic’s own docs treat it as a second call, not a first one. It is built for hour-long autonomous sessions, dense filings and PDFs, research where the model has to follow leads and verify them. It is also twice the price of Opus 5 and, in independent testing by Artificial Analysis, roughly 60% more per task at maximum effort. Would I hire it to investigate something complicated and consequential? Yes. To rewrite a two-sentence email? That is a consulting invoice for a clerical task.
Claude Opus 5 runs the place day to day. It is the documented default, it verifies its own work, it recovers from its own errors, and it scores higher on the independent Intelligence Index than either of OpenAI’s current flagships while costing half of Fable.
Claude Sonnet 5 is the VP who ships. Production volume, automation, multi-step engineering, at $2 in and $10 out per million tokens. This is where most small businesses should be living and mostly are not.
Claude Haiku 4.5 is fast and cheap and has been heads-down since February 2025. That is its knowledge cutoff. Brilliant coordinator, has not read the news in a year and a half.
GPT-6 Astra shipped September 3 as the newly hired CTO with an enormous reputation. Its computer-use score is the best published number in this comparison and it hallucinates about half as often as its predecessor. It is also priced two and a half times higher than Sol while scoring level with it on the independent intelligence index. The coding and computer-use gains look real. The general intelligence gain is not visible in independent testing yet.
GPT-5.6 Sol is the one who goes and finds out. Browsing and multistep research are its strongest published capabilities. Terra is the director of everyday work and the cleanest true match to Sonnet 5 in the whole comparison. Luna is the intake desk at twenty cents per million input tokens.
Two places the comparison breaks, and I would rather say so than force it. OpenAI has no clean equivalent to Fable 5.1, because it escalates by raising effort inside a model rather than by hiring a different one. And Anthropic has no clean equivalent to Sol, because it does not ship a research-specialized model.
The part almost nobody talks about
Here is what changed my thinking more than the model tiers did.
Both companies now let you set how hard the model tries. Anthropic calls it effort and offers five levels. OpenAI calls it reasoning effort, and on the GPT-5.6 family it added an Ultra mode that runs four agents in parallel.
The range inside a single model is bigger than the gap between neighbouring models. Artificial Analysis measured Fable 5.1 producing 13.1 million output tokens at low effort and 143.7 million at maximum. Same model. Eleven times the work.
11x
The work range inside one model
3 pts
What 60% more money buys you
1
Models sitting in the CEO chair
So you are not only choosing who to hire. You are choosing how many hours to give them. A senior person with twenty minutes and a mid-level person with two days produce very different work, and the model name alone does not tell you which one you just bought.
Anthropic went a step further and now lets you change effort mid-conversation on Fable 5.1 and Opus 5, without losing your cached context. Same hire, dialled up for the hard step, dialled back down for the routine ones, inside one engagement. That is not a model choice. That is management.
The escalation policy, as three steps
I have said for years that you test on three before you run seventy-five. This is the same discipline wearing different clothes, and here is the sequence I would hand anyone.
That last clause is the whole thing. Most people reach for the most expensive model because the work matters to them emotionally, which is not the same as the work being hard. Escalate on evidence.
The numbers back the discipline up. On the independent index, Opus 5 at maximum effort scores 63 for about $2.34 a task. Fable 5.1 scores 66 for about $3.76. Three points for sixty percent more money is a real trade-off, and for most work it is the wrong side of it.
What this means if you are running a small business
I do not think the interesting story is that models got smarter. I think it is that we have stopped buying a product and started staffing a bench, and the skill separating people getting real value from people who are not is no longer prompting. It is delegation.
This is genuinely good news if you are a team of one. You do not have a staff. You now have a bench. The constraint has moved from capability to judgment about when to spend, and judgment is something a one-person business can actually manage.
Two failure modes worth naming before you go build on this. The first is paying senior rates for clerical work, which is the common one and the expensive one. The second is quieter: Fable 5 was withdrawn from the market for nearly three weeks in June under export controls. If your entire workflow depends on one top-tier model, you have a single point of failure and you will find out about it on the worst possible morning.
None of this means models replace employees. The selection problem now resembles staffing. The delegation is of tasks, not of jobs. Somebody still has to define the work, judge the output and own the result.
Which is the part of the analogy I find most useful. In this org chart, there is nobody in the CEO chair.
That seat is yours. The whole thing only works if you sit in it.