Content debt: enterprise AI and its failures

Content debt: enterprise AI and its failures

The expected value of AI is still largely… expected. One of the factors, behind this disappointment, is nevertheless little known: it is the content debt.

Hearing the promises surrounding AI, we should have already entered a new era of productivity. By 2024, a BCG study reveals that only 4% of companies are creating substantial value with AI. The following year, Gartner estimated the proportion of AI projects interrupted during 2026 at 60%. Here we are. Indeed, the expected value is still largely… expected. One of the factors, behind this disappointment, is nevertheless little known: it is the content debt.

When the document becomes the problem

AI models are advancing at an impressive and, in some ways, dizzying speed. However, companies are still struggling to transform the test. This discrepancy, eloquent, deserves attention. What if the explanation was not in the algorithms, but in the data we entrust to them?

Consider that when an enterprise AI generates a response, it is not working from a general knowledge base. It will search for the data available through the sum of documents and knowledge made available by the organization. However, a confusing, contradictory or obsolete raw material is of poor quality. Regardless of the generative AI model used, it will not spontaneously reconcile divergent or contradictory information. The most powerful engine remains dependent on the quality of the fuel.

Nothing new under the digital sun. Long before AI, this difficulty was already present. How much time is wasted searching for information? Which document to check the correct version of a process?

AI doesn’t create disorder. She reveals it

Word, PowerPoint, PDF files: we have become accustomed to recording knowledge in office documents. They are often excellent tools when it comes to fulfilling their primary function. Writing, collaborating and disseminating, very good.

But they also have a weakness, not the least of which is that they enclose the content. They compartmentalize it. They bury him. They hinder his mobility. Once encapsulated in its document, the information loses flexibility. Difficult to find it, difficult to reuse it. So, we copy/paste, adapt and duplicate again and again. Cloning information is dangerous. He often makes it diverge from the initial information to the point, sometimes… of contradicting it. By multiplying the same information, we degrade it, we make it lose its solidity, and therefore its value.

Concretely, this is the procedure updated here, but not there. The version that circulates by e-mail, while another is stored on the shared space. The identical duplicates itself, again and again. The single source of truth is lost.

What is a viable knowledge base?

A poorly powered generative AI still attempts to produce coherence. This is what is asked of him. It brings together, aggregates, reformulates and interprets the available information. While anti-hallucination rules limit overt inventions, they cannot resolve contradictions in the knowledge base. When internal sources diverge, poorly supervised corporate AI risks implicitly arbitrating. Not to deceive, but to construct a plausible answer from incoherent documentary material. It then constructs plausibility where we expected truth. Its logic is not that of a proof: it simply seeks a statistically coherent generation.

In other words, by feeding them poorly, we create the conditions that favor uncertain responses. To move towards success, the real question at the origins of the AI ​​project must not be “Which model to choose?”. It would be more: “What quality of knowledge will we entrust to him?” And that changes everything.

We can say that an exploitable knowledge base is based on four key principles:

● Identify the single source of truth. Critical information should only exist once, in one place.

● Define clear governance. Who creates? Who has access? Who validates?

● Enrich content with metadata. They determine its nature, status, scope of use, validation levels and even relationships with other content.

● Make this set sustainable through regular controls and audits.

On paper, this all makes common sense.

In practice, however, these rules come up against a fundamental difficulty: the document itself…

Going towards structuring and the knowledge graph

The traditional document brings together information of very different types. The safety warning, data table or illustration coexist. Without distinction. This document can be validated without all of its components being validated. Its updating is sometimes partial. This is why highly regulated sectors (aeronautics, defense, etc.) have opted for structured content, although only partially (for technical documentation, for example).

It is then a matter of breaking down and organizing knowledge into unique components, each with its own life cycle, metadata and governance rules. Result: every publication results from a dynamic assembly of validated content. This approach is also emblematically illustrated by the DITA XML standard. The benefits are innumerable, starting with the possibility of multi-channel publication, the parallelization of tasks, gains in terms of traceability and even the reduction in translation costs.

We must add another strategic contribution, less often mentioned.

By this articulation of reliable content, by their permanent enrichment with metadata, the company gradually establishes, and almost without thinking about it, a knowledge graph. What she knows takes the form of a living map with multiple entries (by domain, language, product, geographic sector, version, etc.). And this is precisely the type of environment that increases AI performance tenfold. The tool thus responds to queries without having to reconstruct an approximate context (and costly in tokens) and its user has the certainty of work that is based on a single source of truth. Speed ​​of execution adds to reliability.

For a long time, this approach was considered a fad of technical writers in demanding industrial sectors. They worked to ensure that knowledge was governed like other assets, without being heard as it should have been.

The generalization of the uses of artificial intelligence pays tribute to them.

Jake Thompson
Jake Thompson
Growing up in Seattle, I've always been intrigued by the ever-evolving digital landscape and its impacts on our world. With a background in computer science and business from MIT, I've spent the last decade working with tech companies and writing about technological advancements. I'm passionate about uncovering how innovation and digitalization are reshaping industries, and I feel privileged to share these insights through MeshedSociety.com.

Leave a Comment