After large model service providers gradually shifted from a subscription-based model to billing based on token, the internal usage costs of AI within enterprises began to rise rapidly. McKinsey stated that the company is not currently imposing rigid limits to reduce employee usage, but rather first allows high-frequency users to clearly see their own consumption patterns, and then combines this with technical means to improve efficiency.
10% of users consumed 65% of the usage.
The person in charge of the AI business under McKinsey, Patnaik, stated that the company compares this kind of reminder mechanism to a mobile phone data usage alert. The focus is not on preventing employees from using it, but on telling them how to complete the same work at a lower cost.
According to the company's blog data, as of May 2026, McKinsey processes approximately 5 trillion AI token per month. The distribution of usage is not even; about 10% of users account for 65% of the total consumption. Among them, consultants and software engineers are the main user groups.
- Monthly processing volume is approximately 5 trillion token.
- Approximately 10% of users consume 65% of the total amount.
- The main groups with high usage are consultants and engineers.
Internal gateway and circuit breaking mechanism have been launched.
Patnaik says that McKinsey's current AI expenditures have not yet reached an out-of-control stage, as most of the use cases are still focused on improving individual efficiency and delivering customer projects more quickly. The benefits of such investments are still temporarily higher than the costs.
However, the company has implemented multiple layers of control measures. Its internal AI gateway optimizes the requests before sending them to the model service providers; when there is abnormally high token consumption in certain scenarios, the system will temporarily suspend access to first determine whether such usage truly generates value.
In addition, McKinsey also uses a caching mechanism to reuse answers to repeated questions, thereby reducing the need for repeated calls. The company also manages the consumption of different business units uniformly, rather than locking unused quotas in separate licenses.
The consulting industry has begun to focus on reducing token expenditures.
Similar practices have also appeared in other consulting firms. Ernst & Young stated that they have established a dedicated team to manage the investment in AI, and have deployed a routing system at the backend of some professional tools to assign employee requests to models that are more suitable for the tasks. According to Ernst & Young, since April, this system, in conjunction with governance measures, has reduced token consumption by 60%.
A senior software engineer at Deloitte in the United States stated in June of this year that after GitHub adjusted its billing method, developers would quickly use up their new monthly quotas, and as a result, the team's expectations for the work pace were also affected.

Patnaik It is also indicated that McKinsey officially established a related business this spring to help clients deploy AI at lower costs. Their approach is not to simply look at how much token each employee uses, but rather to consider the cost required to achieve a result. If one only pursues fewer token, it may instead suppress the use of valuable AI.












