Last week, a managing partner told me some version of the same thing I keep hearing. He is exhausted. Every few days there is a new flagship model, a new agent story, a new price drop, a new warning about lock-in, a new demo that looks magical for ten minutes and forgetful by lunch. He is not confused because he is behind. He is confused because the market keeps shoving the wrong decision in front of him.
The wrong decision is, which model should we bet on?
For a mid-sized law firm, CPA firm, brokerage, executive search firm, or advisory shop, that is increasingly the least important strategic question.
Look at what just happened. Google’s research systems are doing things that feel almost science fiction at first read, an agent found a new way to multiply 4x4 matrices in 48 scalar multiplications, beating a record that stood for 56 years. The same system reportedly recovered 0.7% of Google’s worldwide compute for over a year, and another fix cut Gemini training time by 1%. One autonomous research system ran for 417 hours and produced 166 fully AI-generated papers for about $180K. That is real capability. But sitting right next to it is an equally important business signal, Grok 4.5 launched at $2 per million input tokens and $6 per million output tokens, with one cited benchmark putting it at just $0.31 per task.
That is the story I think leaders are missing. Intelligence is getting better, yes. It is also getting cheaper, faster, and easier to embed.
That changes the economics for professional-services firms in a very specific way. If frontier-grade intelligence keeps sliding toward commodity pricing, then buying access to a smart model stops being a moat. It becomes table stakes. Simon Willison’s rundown of the GPT-5.6 family makes that plain, three sizes, Luna at $1/$6, Terra at $2.50/$15, Sol at $5/$30 per million input/output tokens
For most firms in my Goldilocks Zone, 30–500 people, this is actually good news. You do not need to win the model wars. You need to stop overpaying emotionally for them.
The second signal matters just as much. The major labs are climbing into the application layer. OpenAI is openly calling ChatGPT the start of a “superapp.” Meta’s Andrew Bosworth says the “big monolithic model era is over” and that value will come from how models are connected to product. Narayanan and Kapoor warn that this move up the stack helps labs escape the commodity trap, but it raises the risk of customer lock-in.
That sounds abstract until you put it in the shoes of a 120-person accounting firm or a 75-lawyer firm. If your workflows, prompts, client knowledge, review habits, and staff behavior all get wrapped inside one vendor’s environment, you may get speed now and dependency later. I think that is the practical implication. Not some distant policy debate. You have to build the ability for model-agnostic behavior into all of your AI solutions.
And this is where I come back to my 10-20-70 Rule. Technology is only 10% of a successful AI transformation. The market keeps shouting about the 10. Your returns will come from the 70, people, trust, training, review discipline, and workflow redesign, plus the 20, your process and data rework. If model quality keeps converging and price keeps falling, the human side becomes even more decisive, not less.
This is also why I frame the answer through my Centaur Firms model. The firms that pull ahead will not be the ones with the flashiest model subscriptions. They will be the ones that pair lower-cost intelligence with higher-value human judgment inside the work itself. In practice, that means a valuation team that uses AI to prepare the first analysis pass, but keeps partner attention on exceptions, interpretation, and client implications. It means recruiters using AI to draft prep briefs and long lists, while consultants spend their time reading the human beneath the résumé. It means a CRE team generating submarket drafts in minutes, then using live market instinct to challenge and sharpen them. These will be companies that look at processes across the entire company, not just fixing a patch on a single process.
The model is the electricity. Your workflow is the machine.
A managing partner should make one concrete decision this quarter: pick one core workflow where your professionals are still spending too much time assembling, summarizing, drafting, or hunting, and redesign it as a Centaur workflow with model independence, clear human approval points, measurable time savings, and client-safe governance. Do that before you sign a bigger enterprise contract. Do that before you let vendor convenience become firm dependency.
The intelligence is getting cheaper every week. Whether your judgment becomes more valuable because of it is up to you.
The rest of this brief examines how the conversation should be opened — what specifically to say to clients who are silent, how to structure the disclosure so it lands as competence rather than panic, and three ways the Friction Audit reveals exactly which client conversations to have first…