The Real Reason Your Business Is Burning Too Many Tokens
Kognitos Founder & CEO Binny Gill holds nearly 100 patents in computer science.
gettyThe biggest worry I hear from enterprise leaders right now is the cost of AI. Teams are consuming tokens faster than their budgets anticipated, and the value coming back rarely feels commensurate. The instinct is to blame the technology or wait for prices to fall. I believe the problem is older and more familiar. Forget technology for a moment and go back 100 years.
A business a century ago hired many kinds of people. A few were brought in for strategic judgment—perhaps the chief finance officer or chief legal officer. They were paid well because that intelligence was expensive and mistakes carried serious consequences.
Factory workers were interviewed differently. Nobody asked whether they could strategize. They were asked whether they could learn the machine, follow the rules and repeat the work accurately. They cost less, and that was the point. You could put someone with the mind of a vice president on the assembly line, but it would waste money.
Every human is intelligent in different ways. AI comes in the same gradations. Some models are built for strategic, creative reasoning, whereas others understand instructions and carry them out. We know how to manage this because we’ve been doing it for centuries.
Businesses also learned to reject candidates for being overqualified. The concern wasn’t only that they might leave. Someone too creative for a rules-based role may question the process and introduce changes where the business simply needs the rules followed. Too many cooks is a problem.
If you hire the smartest person to carry medicines between hospital floors, one day, they may devise a human relay so they no longer have to walk. Then, the medicines get mixed up. That’s what an overly intelligent AI can end up doing—it tries to be smart about a job that only required following the steps. Although smartness creates efficiency, it usually comes with increased risks.
Getting a smarter AI system doesn’t mean getting a safer one. It depends on the job.
A peer-reviewed 2026 study published in Organization Science, involving 758 consultants, found AI improved performance on tasks it could handle reliably. On a task outside those capabilities, people using AI were 19% less likely to reach the correct answer than those working without it. The AI helped when the work suited it and made performance worse when it didn’t.
We’ve always reduced error by reducing the intelligence deployed to execute a task. That’s why the calculator exists. Once math is written down as rules, a calculator can execute it without errors, unlike humans who are smarter but more error prone.
Humans figured out how to get many people to do the same work consistently. Explaining the same task to 10 people is expensive, so we write the manual once. The strategic thinking happens up front. From then on, we need people who can read plain English and follow step one to three.
That’s how a coffee shop makes the same latte every time, and how a fast-food chain makes its fries taste identical no matter which high schooler is on shift. Writing the steps down lowers the cost of the intelligence you need to hire because it takes a higher intelligence to plan and lower intelligence to execute the plan.
If your business is burning too many tokens, it’s making one or both of two mistakes. The first is using AI that’s overqualified for the work. The second is handing AI work that hasn’t been documented properly. The model is forced to reason a lot more each time than it should.
Spend the tokens once to write the rules down, then use a cheaper form of intelligence to execute them.
Stanford researchers tested this approach with FrugalGPT, which chooses among different language models based on the query. In their experiments, it matched the performance of the best individual model while reducing costs by up to 98%. The exact savings will differ by workload, but businesses clearly don’t need to use the most capable model for every job. I posit that cheaper, specialized models will gain popularity in the coming months, as they’ll produce less bias and higher accuracy at constrained tasks.
Some enterprises are already working this out. A CIO at a large financial services firm told me recently that he no longer gives AI to everyone because most people don’t yet know how to use it carefully. Instead, he built a layer on top of the models, assigning the right intelligence to the right agent persona for the right job. Then, the employees were given access to the AI agents. This is AI rationing, and I think we’ll see much more of it until the managers build an appreciation for tokenomics.
What’s missing is the instinct managers already have with people. Any hiring manager knows when a role needs an intern and when it needs a specialist with 10 years of experience. Nobody yet reads AI that way. AI vendors haven’t helped. They keep presenting the newest model as the better model, without teaching businesses that an older or smaller one can do the job.
Imagine if every model came with a resume. This one is a high school graduate. This one holds a Ph.D. This one has a decade in your field. You could hold your job description next to it, interview the model and hire and pay accordingly.
The answer to the cost of AI isn’t simply cheaper AI. Businesses already know how to match different levels of human intelligence to different jobs. They now need to do the same with AI: Choose the right model for the work, write the rules down once and reserve expensive reasoning for tasks that genuinely require it.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


