Autonomous agents reveal the cost of poorly prepared data

Autonomous agents reveal the cost of poorly prepared data

The rise of autonomous agents capable of querying databases or generating code to automate analyzes is transforming the way organizations use their data.

Since the rise of generative AI, these agents have attracted growing interest among companies, which see them as a way to accelerate the analysis and exploitation of data. However, behind this promise of simplification lies a reality that is often underestimated. To be truly effective, these agents must rely on reliable, clearly structured and well-documented quality data. A prerequisite that many companies still struggle to guarantee.

AI challenged by data complexity

In many modern data environments, information remains fragmented between different data warehouses, technical pipelines and analytical tools. The data is accessible, but its interpretation often remains uncertain. The same metric, whether revenue or number of active users, can vary depending on how it is defined or queried, as well as the teams using it. These discrepancies pose challenges for humans, who are nevertheless able to identify and correct them thanks to their implicit knowledge of the business context. For an AI, on the other hand, they immediately constitute a breaking point.

An agent may produce a technically correct SQL query while relying on an inaccurate definition of a key metric. The result then seems coherent, but it can in reality fuel biased or misleading analyses. The more companies seek to automate their analyzes using agents, the more the limits of their data foundations come to light.

The hidden cost of data preparation

These difficulties are not just technical. They also weigh, in a very concrete way, on team productivity. Today, a large part of analytical work remains devoted to preparing and verifying data before it can even be used. According to a study of data professionals, analysts spend an average of 9.1 hours per week, or more than a day of work.

At the enterprise level, this inefficiency creates a significant cost. It slows down the production of analyses, delays decision-making and limits the effective impact of data teams. Autonomous agents relying on inaccurate data only exacerbates this reality.

Structuring data for agents and humans

To take full advantage of AI, organizations can’t just make their data accessible. They must also clarify their meaning and rules of interpretation. This is precisely where the semantic layer comes in, which connects raw data to business uses by providing a common language, shared metrics and explicit interpretation rules, essential for reliable use by AI systems. In such an environment, each metric benefits from a unique definition, consistent from one tool to another and common to all teams. Dashboards, analytical applications and AI agents then rest on the same interpretation base. This consistency improves the reliability of analyzes and limits ambiguities, while providing automated systems with a more solid basis for exploiting the data.

This is what Anthropic found, following their adoption of analytical agents: without a good metadata foundation, they could not have obtained more than 21% accuracy (rate of correct answers to analytical questions) because of the ambiguity of the data, its lack of freshness and its discoverability. When they were able to run the same model on a better governed and documented data structure, efficiency increased to 95%. Even though they don’t claim to have used a semantic layer, their explanation reads like an enumeration of dbt’s features.

The rise of autonomous agents does not only modify the methods of producing analysis. This phenomenon also increases the requirements to which data architectures are now subject. The organizations that will be able to take full advantage of these new technologies will be those that have invested in the structuring and governance of their data. Others may find that automation not only speeds up analysis, but also the propagation of errors that are already present. In this context, the question is no longer just how to integrate AI into analytical workflows. It consists above all of guaranteeing that the data on which these systems are based are understandable and reliable but also correctly contextualized. Autonomous agents can only produce relevant results if they have a clear framework: shared definitions, explicit business logic and validation mechanisms to correct erroneous interpretations. As autonomous agents are deployed, data quality ceases to be a purely technical subject and becomes a strategic imperative. It is not only a question of making data reliable but also of giving systems the means to interpret them correctly: two essential conditions for stimulating productivity and bringing new perspectives to the surface.

Jake Thompson
Jake Thompson
Growing up in Seattle, I've always been intrigued by the ever-evolving digital landscape and its impacts on our world. With a background in computer science and business from MIT, I've spent the last decade working with tech companies and writing about technological advancements. I'm passionate about uncovering how innovation and digitalization are reshaping industries, and I feel privileged to share these insights through MeshedSociety.com.

Leave a Comment