Organizations have spent the last few years looking for answers to fundamental questions about AI: Is AI effective? Can it improve productivity? What organizational benefits does AI provide?
Now, with those questions largely answered, enterprises face a new question that will ultimately determine their ability to grow their businesses and remain competitive: How can they scale AI for 20,000 or more people each day, at a cost the organization can support?
The hidden cost of AI at scale
Traditional budgeting methods no longer work in today’s AI economy. As quickly as AI has accelerated, so have the hidden costs of the technology. Most businesses budgeted for a known fixed-seat cost — and were surprised by high invoices when their vendors moved to a usage model. With AI, and agentic AI in particular, it can be harder for organizations to determine what they are actually paying per answer.
“Agentic, retrieval-intensive workflows can require a dozen or even 20 model interactions to complete a single task,” says Jason Banta, vice president of global go-to-market compute sales at Qualcomm Technologies, Inc. “Many IT leaders are not accounting for this type of workflow and number of calls to the models in their budgets.”
Each time an AI model processes information, the model breaks down the text or image into tokens, which many vendors now use as the metric for billing AI usage. Gartner reported in August 2026 that while agentic models perform many more tasks than a human using GenAI, they require 5 to 30 times as many tokens as a standard GenAI chatbot.
What’s more, Banta says, the cost challenge will continue to grow. It’s tempting to think that falling token prices will help with costs, especially since the Stanford 2025 AI Index found that the cost of querying a model performing at roughly the level of GPT-3.5 decreased from $20 to $0.07 per million tokens between November 2022 and October 2024. But even with lower token costs, Gartner predicts that higher token consumption will raise overall inference costs.
As a result, leaders are changing how they think about their AI budgets, with tokenomics becoming an increasingly important part of their AI cost strategy.
Rethinking AI costs through tokenomics
As organizations started exceeding their AI budgets only a few months into 2026, many immediately targeted employee usage by limiting tokens, a move that quickly proved problematic. This type of blanket approach may limit employees not only from using excessive tokens to create an internal presentation but also from using AI to help identify new business opportunities that could be worth hundreds of thousands of dollars in revenue.
The pace of change requires organizations to adapt to real-time learning and evolve with both technology and new use cases. Banta says a demand-based approach helps organizations best determine where to deploy each capability based on the value it returns. Leaders must also address key infrastructure decisions, including how AI workloads are orchestrated.
“A lot of organizations are building the plane in the air,” Banta says. “Many are in a discovery phase with orchestration between small models, medium models, and large models, as well as orchestration between on-device, on-prem, and cloud.”
The on-device shift: From cloud dependency to intelligent architecture
As smaller and midsize models have significantly improved accuracy and performance, enterprises are rethinking the assumption that AI workloads should default to the cloud. Gartner recommends that organizations use small, domain-specific language models for routine, high-frequency tasks, which typically perform better at a fraction of the cost, especially when aligned with specialty workflows. The research also recommends organizations heavily gate expensive frontier-level models and reserve these for high-margin, complex reasoning tasks.
Organizations wanting to reduce inference costs, energy consumption, and their overall AI carbon footprint are increasingly turning to using on-device AI for these models, Banta says, and avoiding the cloud altogether. He says that high-frequency AI workloads such as summarization, transcription, drafting, translation, search, and image cleanup can run effectively on smaller or midsize models directly on a phone or AI-ready PC powered by capable silicon, such as Snapdragon®.
In addition to reducing token costs, on-device AI addresses the increasing concerns about AI and energy consumption. The U.S. Data Center Energy Usage Report estimates data centers could account for 9.5% to 15.3% of total electricity consumption in the US by 2030.
“With energy availability, grid interconnect, and standing up power, power energy capacity is becoming scarcer, which creates bottlenecks in deployment,” Banta says. “We found that a green inference is one that doesn’t leave the device. The Snapdragon platforms have a massive amount of capacity to handle the smaller and medium-sized models that answer key questions.”
On-device AI also addresses the challenges of managing and classifying data in the AI era, especially in regulated industries such as finance and healthcare. With on-device AI, data stays on the device, providing another layer of security at the endpoint. Because agentic AI requires maintaining distinct control over agentic behavior, on-device AI for agents can provide a higher level of both control and privacy, Banta says.
TCO beyond the sticker price
Scaling AI requires a solid tokenomics strategy that evaluates the model and compute environment for each workload to ensure the enterprise gets sufficient value from the tokens consumed. Enterprises should evaluate the total cost of AI across the entire stack, Banta says, including on-device, on-premise infrastructure, and cloud, over the lifecycle of an employee’s AI usage.
“Leaders need to use a completely new type of math for calculating TCO in the AI era,” he says. “Traditional metrics such as device costs and seat licenses tell only part of the story. To deploy AI effectively at scale, organizations need a holistic understanding of annual per-user costs, including infrastructure, model usage, and operational expenses.”
While tokens make it very easy for organizations to go over budget on AI costs, you pay for endpoint purchases only once. Organizations then gain benefits for the entire device lifecycle, which averages four years for laptops, according to Sage’s 2026 IT Asset Management Benchmarking report. Instead of focusing on sticker prices, enterprises should consider investing in AI-capable endpoints, which can offset high and recurring inference costs in as little as three to five years, Banta says.
Building an AI architecture for sustainable scale
Because every enterprise is unique, with different usage patterns, employee personas, and workloads, a universal formula for AI architecture does not exist, Banta says. Instead, each business must build its own framework to deploy AI at scale.
Many business leaders are making AI-related decisions that affect their companies’ futures, with the goal of simply reducing costs. Organizations that instead focus on creating an architecture and process that delivers value at a sustainable cost can achieve what few others are doing — affordably deploying AI at scale.
“When we work with enterprises to design deployment for their organization, we look carefully at where models are deployed for many routine tasks in agentic workflows, such as classification and summarization, which is the majority of how people utilize AI today,” Banta says. “We determine that many of these workflows can be done directly on-device with smaller and mid-sized models. The ability to place AI workloads where they make the most sense, on-device or in the cloud, starts with a foundational choice: the silicon platform powering an organization’s AI strategy.
“As a solution provider for both hardware and software, Qualcomm Technologies provides an array of solutions across device categories, including the edge, and data centers that answer many of the questions being asked in boardrooms,” Banta says. “Qualcomm’s portfolio is designed to support that flexibility, helping organizations optimize for cost, performance, security, and user experience simultaneously. The result is an AI strategy built not just for experimentation, but for sustainable enterprise-wide adoption.”
Discover how Snapdragon can help your organization get the most business value from on-device AI.
This post was created by BI Studios and sponsored by Qualcomm Technologies, Inc. and/or its subsidiaries.
