Ernst & Young stated that in April of this year, the company launched an internal AI routing system globally to distribute employee requests among different models. The way employees use it remains largely unchanged, but the system in the background determines which model is more suitable for the current task, in order to control call costs and improve response efficiency.
Background allocation of model requests
This system is referred to by EY as an "invisible router." It currently integrates with some specialized AI tools, primarily covering scenarios in departments such as taxation and risk management. It is not fully utilized in all internal products, nor is it suitable for the Microsoft Copilot chat tool, which is open to all employees.
According to EY, employees still input prompts and retrieve results as usual, but the backend directs requests to different models based on task complexity. For tasks that do not require high-level inference, the system avoids directly calling larger, more expensive models.
Dan Diasio, EY's Global AI Lead, explained that heavier models typically require more inference processes, thus taking longer to return results. The role of routing mechanisms is to reserve higher-cost models for scenarios where they are more needed.
Token-based billing helps companies control costs.
As more AI service providers adopt token-based billing, enterprises are paying significantly more attention to the cost of using models. The article mentions that from February to June of this year, OpenAI, Anthropic, and GitHub all introduced pricing methods linked to token usage; if enterprises use high-performance models for simple tasks, the cost may increase rapidly.
EY cited its AI Pulse survey released this week in the US, stating that among 534 senior decision-makers in US companies, 82% of respondents expressed concern about token usage. The survey was conducted from April to May.
Dan Diasio believes that many companies suddenly incur high bills not necessarily due to widespread use, but rather because a few employees are consistently using the wrong tools. The goal of routing systems is to reduce this kind of waste.
Enterprise AI management shifts to the budget stage
Ernst & Young stated that after the router was implemented along with training and governance measures, token consumption decreased by up to 60% since its launch in April. The company also set token budgets by position and department, and employees exceeding their budgets need approval to apply for more usage.
In the same survey, EY also stated that 64% of the executives surveyed said their organizations had begun monitoring AI Token usage and had set clear budget constraints. Diasio expects this percentage to continue to rise in the coming months.

He also stated that in the next phase, companies will not only track how much AI employees use, but will also pay more attention to whether the investment brings actual output, including whether it shortens task completion time and whether it motivates employees to take on new work content.












