Control AI costs before they explode

Control AI costs before they explode

As AI becomes more widespread, its costs become difficult to control. Find out how to better track usage, allocate expenses and choose the right model for each need.

Generative AI has quickly moved from the stage of isolated pilot projects to that of an infrastructure widely used cross-functionally within the company. The first feedback shows, however, that this increase in power is accompanied by a major challenge: cost control. Uber, for example, announced that it had consumed its entire annual AI budget in just four months, while acknowledging that it was not yet able to establish a link between this expenditure and the productivity gains obtained.

As use cases multiply, model consumption increases, as do the associated costs. Teams often use vendor models via a shared API key, with billing by account level, expressed in tokens. In this context, it is impossible to know which user, team or automated process is actually generating the costs. Therefore, effective control of expenditure requires being able to precisely link consumption to its origin.

This lack of visibility is not without consequences. The financial implications become clear when comparing AI spending to other budget items. HR or licensing costs are generally allocated according to the annual plan, which makes it possible to evaluate them and control their evolution. For AI consumption, this ability to control and manage costs is almost non-existent. Without a clear view of actual spending, it is impossible to measure the contribution of an AI investment. And without attribution to the teams and processes involved, it is just as impossible to limit spending in a targeted manner.

Even more problematic, in the absence of a defined budget, transparency, or clear criteria for choosing a model, users tend to gravitate toward the best-performing model available, even when a simpler, less expensive alternative would deliver exactly the same result. This problem is unfortunately likely to last. If Gartner estimates that the cost of inference could fall by 90% by 2030, companies’ bills should not decrease by the same proportion. The reason? Ever more intensive uses and the rise of film models.

Go beyond the limits of traditional mechanisms with a central control point

To control these uncontrolled expenses, companies might be tempted to set a quota of requests. However, while this makes it possible to cap requests and avoid load peaks, their number is only a very rough indicator of the real cost.

Indeed, not all models display the same price and, depending on the context, costs can vary considerably. A single query to a high-end model, with a large volume of data, can cost more than hundreds of queries to a compact model. Even invoicing in tokens remains difficult to use for managers. Indeed, tokens do not constitute an intuitive control unit. Solving this costly puzzle requires a mechanism that can measure spending in monetary value and control it independently of sheer query volume.

An effective approach consists of intervening at the point where requests leave the company’s IS. When access to different model providers goes through a centralized middle layer, called the AI ​​Gateway, all requests go through a single point of control. This gateway makes it possible to monitor consumption and costs, regardless of the supplier used, thus offering a complete vision of expenses and making it possible to think simply in euros (or dollars), rather than in tokens.

Once this visibility is obtained, budgetary control must be granular. Spending limits must be able to be defined by model and/or supplier, but also according to different criteria such as user or application. Budgets can thus be set according to needs, daily, weekly or monthly; For each request, the estimated cost is calculated from the model price and compared to the available budget. Once the cap is reached, new requests are either blocked or redirected to a less expensive model to avoid any service interruption.

Assign uses and costs to the right people

For budgets to be truly actionable, you need to know the consumption of each service, or even of each employee. If the company is satisfied with the information transmitted by the applications themselves, the monitoring remains unreliable. The attribution must be done directly at the gateway level, based on a verified identity. When a user logs in through the internal authentication tool, their identity is securely associated with each request. It will therefore be possible to track consumption by user, team or application, and allocate costs with much more precision. On this basis, some teams may be allowed to use more powerful models, while others will use lighter, less expensive models.

But taking advantage of AI and optimizing it requires choosing the right model for the right use. The goal is not just to spend less, but to make better use of the models available. Indeed, not all tasks require the most powerful model: a summary can be made by a lighter model, without loss of quality, when a complex analysis can justify the use of a more efficient model. This is where intelligent routing comes in because it allows you to analyze a request and direct it towards the most suitable model, offering the best balance between quality, performance and cost. Management no longer consists only of capping expenses, but of optimizing them.

Beyond the financial aspects, this subject goes beyond just the budgetary question and also touches on governance, data protection, and in certain countries like Germany, social dialogue. Before setting up individualized monitoring, the teams responsible for data protection and, where applicable, staff representatives must be involved to understand actual uses and adapt internal usage limits. Conversely, if the latter are too strict, they can cause use towards uncontrolled channels.

Better control to better promote AI

With the widespread use of generative AI, the question is no longer just how to use it, but how to control its costs and measure its value. In practice, as long as consumption cannot be linked to its origin, AI remains an expense that is difficult to control. To regain control, companies must have a central point of passage that makes uses and costs visible; budgets expressed in monetary value rather than tokens and reliable attribution based on the identity of the company’s users.

Combined with a selection of models adapted to each task, this approach not only limits expenses, but also optimizes them. AI then becomes a resource that is monitored, budgeted and managed like other company resources.

Jake Thompson
Jake Thompson
Growing up in Seattle, I've always been intrigued by the ever-evolving digital landscape and its impacts on our world. With a background in computer science and business from MIT, I've spent the last decade working with tech companies and writing about technological advancements. I'm passionate about uncovering how innovation and digitalization are reshaping industries, and I feel privileged to share these insights through MeshedSociety.com.

Leave a Comment