Beneath Tokenomics: Why $/FLOP Is Becoming A Critical AI Cost Metric

Direct Source Verification: This story is aggregated from Forbes (forbes.com). Full reporting rights and copyright belong to the primary publisher.
Tech leaders need to understand how efficiently capital, energy and physical capacity are being converted into computation.

Vinod Bijlani is an AI practice leader at Hewlett Packard Enterprise.

gettyOne number comes up in almost every cost discussion I have with CIOs and AI leaders: price per million tokens.

This metric is useful for comparing inference services and forecasting application spend, but it only measures the price of consuming AI, not the productivity of the infrastructure producing it.

That distinction matters as AI usage shifts from occasional prompts to persistent agents, multimodal workflows and reasoning-intensive applications. A low token price can conceal underutilized accelerators, memory bottlenecks, network congestion and rising power costs.

Leaders, therefore, need to ask a second question: How efficiently are capital, energy and physical capacity being converted into computation?

This is where dollars per floating-point operation ($/FLOP) becomes strategically useful. A token is what an application consumes. A FLOP is one unit of the computation the infrastructure performs.

Tokenomics can help explain consumption economics, and datanomics addresses the economics of making data reliable and usable, as I have written about previously. $/FLOP adds the infrastructure lens without diminishing the other focuses, because each metric measures a different layer of the AI stack.

Epoch AI’s recent trend data indicates that the performance purchased per dollar of AI chip spending has improved by about 49% a year since 2023, which is encouraging. The organization’s research into algorithmic progress in language-model pretraining also found in 2024 that the compute required to reach a given performance level historically halved roughly every eight months.

Demand, however, is growing even faster. Another piece of Epoch AI estimates that training compute for frontier language models has expanded by about five times a year since 2020.

The precise trajectory will change, but the direction is clear: Cheaper computation is enabling more computation.

This is the paradox at the heart of AI infrastructure economics: Efficiency can improve while total spending rises. However, lower unit costs can make longer contexts, more model calls, richer modalities and larger agent loops economically possible.

Falling $/FLOP doesn’t guarantee a falling AI bill, and it can’t tell leaders whether demand is being controlled. But it does show whether the infrastructure layer is becoming more productive.

Discussions of $/FLOP may be treated as a chip comparison, but enterprise AI is a system-level workload. Its effective infrastructure cost includes accelerators, memory, networking, storage, power, cooling, floor space, platform software and the engineering required to keep the environment productive.

Power, for example, is becoming one of the hardest constraints. The International Energy Agency projects that global data center electricity consumption will roughly double from 485 terawatt-hours in 2025 to around 950 terawatt-hours by 2030. Electricity consumption from AI-focused data centers is expected to grow even faster, roughly tripling over the same period.

Understanding how many ​accelerators you can acquire is, therefore, only part of the equation, while the more pressing question is: “How much reliable computation can we deliver within the power, cooling and space actually available?”​

Meanwhile, published peak performance is not the same as delivered performance. Model architecture, numerical precision, batch size, memory bandwidth, communication overhead and utilization all determine how much installed capacity becomes productive work.

That’s why $/FLOP should not be used as an isolated purchasing score. A more useful enterprise measure is the effective cost per delivered FLOP for a defined workload and service level. This keeps the metric focused on infrastructure productivity: how much usable computation the environment delivers after accounting for utilization, bottlenecks, energy and operational overhead.

This boundary is important to understand. Poor data quality belongs primarily to datanomics. Excessive prompts, context growth and model calls belong primarily to tokenomics. Idle accelerators, memory constraints, network contention and energy inefficiency belong to $/FLOP.

The metrics should connect, but collapsing them into one number hides the decisions leaders need to make.

1. Match the model to the task. Smaller or specialized models can serve predictable, high-volume workloads, reserving more capable models for complex reasoning. This reduces compute demand before new capacity is purchased.

2. Optimize memory and data movement. Many inference workloads are constrained by memory bandwidth, caching and interconnect performance rather than arithmetic throughput alone. Research into large-batch LLM inference has shown how DRAM bandwidth saturation can leave GPU compute capacity underutilized. Improving data movement can unlock capacity that has already been paid for.

3. Raise utilization without weakening service levels. Scheduling, batching, workload placement and capacity sharing can spread fixed infrastructure costs across more delivered computation, provided latency and reliability remain within target.

4. Design power and cooling with compute. Rack density, thermal design and energy availability should be architecture inputs from the beginning. Uptime Institute’s analysis of AI cooling shows why the cooling method increasingly depends on rack power, with liquid cooling typically entering the design at higher densities.

5. Measure the full platform. Accelerator telemetry alone will not reveal storage delays, network contention or orchestration overhead. Connect token cost, latency and calls per outcome with effective $/FLOP, utilization and energy efficiency.

Taken together, these measures can expose trade-offs: A more capable model may cost more per call but complete a workflow in fewer steps, while higher utilization may weaken latency.

Otherwise, optimizing one layer can simply move cost elsewhere in the AI stack.​

Tokenomics gave leaders a practical way to discuss the cost of consuming AI. Datanomics exposes the investment required to make enterprise data ready for it. $/FLOP brings the underlying infrastructure into the same economic conversation.

This metric will not replace token economics, and it should not become another isolated benchmark. Its value is more focused: showing how efficiently an AI environment converts capital, power and physical capacity into delivered computation under real operating constraints.

By understanding the economics of each layer and measuring the trade-offs between them, organizations can continuously improve the productivity of the infrastructure beneath every token.​​​​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Original Source
https://www.forbes.com/councils/forbestechcouncil/2026/09/17/beneath-tokenomics-why-flop-is-becoming-a-critical-ai-cost-metric/
Visit Forbes ↗
SHARE STORY:
𝕏 f in

Related Coverage in Business