As enterprise adoption of generative AI shifts from flat-rate subscriptions to consumption-based token pricing, professional services firms are facing operating expenses. In response, McKinsey & Company and rival firms like EY and Deloitte US are deploying automated usage alerts, internal routing gateways, and circuit breakers to curb escalating compute costs.
The Bottom Line
- The Shift in Pricing: Large Language Model (LLM) providers have moved from subscriptions to consumption-based pricing models tied to token counts.
- Concentrated Consumption: Data from May 2026 indicates that roughly 10% of McKinsey users account for approximately 65% of the firm’s total monthly consumption of five trillion tokens.
- Mitigation Controls: Enterprises are deploying automated gateways, invisible routers, and real-time usage alerts—yielding measurable efficiency gains, such as a 60% reduction in token usage reported by EY.
The End of Freewheeling AI Token Consumption
The economic reality of running large-scale artificial intelligence operations caught up with corporate balance sheets. Earlier this year, major Large Language Model (LLM) providers moved from subscriptions to consumption-based pricing models. Under this new regime, organizations pay based on the number of tokens—individual chunks of text read and generated by the model—driving up overhead as usage grows.
The financial impact is sharp. In September, OpenAI announced that its most prolific users of AI coding agents now consume more than $7,000 worth of tokens daily. For professional services giants, managing spending pressure requires active intervention against runaway compute consumption.
At McKinsey, internal AI traffic hit about five trillion tokens per month by May 2026. According to company disclosures, usage is heavily skewed: about 10% of the firm’s user base generates roughly 65% of total volume, driven largely by software engineers and consultants.
Internal Controls and the Automated Alert System
Rather than capping employee usage, McKinsey opted for transparency and education. Led by Debasish Patnaik, who leads QuantumBlack in the UK, the firm introduced a firmwide alert system over the summer. When an employee’s token consumption gets too high, automated emails notify them.
“Similar to using mobile data on a work phone, we tell them this is how you could do things to make it more cost-effective for the firm,” Patnaik explained regarding the strategy.
Beyond simple notifications, McKinsey relies on an internal AI gateway. This system optimizes requests before they reach external model providers. It also utilizes circuit breakers to temporarily pause access around particularly high token usage while the firm assesses whether the usage is productive.
| Enterprise Firm | Control Mechanism Implemented | Reported Metric / Impact |
|---|---|---|
| McKinsey & Company | Automated user email alerts, internal gateways, and circuit breakers. | Processing ~5 trillion tokens/month (10% of users account for 65% of volume). |
| EY | “AI Value Realization Office” and an invisible task-routing model. | Achieved a 60% reduction in token consumption since April. |
| Deloitte US | Managing GitHub pricing adjustments and monthly developer quotas. | Developers quickly exhausting monthly quotas due to model pricing shifts. |
Competitor Responses and Structural Optimization
The challenge is not unique to McKinsey. Across the consulting sector, firms are pursuing controls to manage AI spending. EY established a dedicated “AI Value Realization Office” and deployed an invisible router behind specialized tools. This architecture automatically steers employee queries to the model best suited for the task. According to EY, those governance measures successfully drove down token consumption by 60% since April.
Meanwhile, software engineering teams face immediate friction. In June, a senior software engineer at Deloitte US noted that recent pricing changes implemented by GitHub are “already wreaking havoc” on expectations for work. Developers rapidly burned through newly imposed monthly quotas that took effect that month.
Advising Corporate Clients on ROI
For Patnaik and the QuantumBlack division, managing internal spend serves as a testing ground for external client advisory. Organizations are grappling with a “hidden cost that we haven’t completely thought through.” Executives are forced to balance AI’s potential competitive benefits against its return on investment.
Patnaik advises corporate clients to evaluate spending based on token cost per outcome rather than per employee. Simply optimizing for fewer tokens could discourage valuable use of AI. For now, maintaining autonomy while enforcing transparent education remains a primary approach to managing AI expenditures.
Disclaimer: The information provided in this article is for educational and informational purposes only and does not constitute financial advice.