A major traditional pharmaceutical company implemented AI by digitizing first and adding intelligence second — and that order is the whole lesson.

This is for leaders whose AI programme is stalling on data nobody can actually read. It is a more common failure than the pilot-to-production gap, and it is much less often named, because the diagnosis is unglamorous.

Trapped knowledge assets

Pharma giants sit on decades of drug research. It is arguably the most valuable asset they own, and most of it is effectively inaccessible: PDF scans, siloed databases, paper documents. Overwhelmingly unstructured data, in disconnected systems.

The consequence is stranger than it sounds when stated plainly: their own scientists cannot use their own archives. Full-text search is slow. Scanned documents are invisible to search engines — a scan is a picture of text, not text. Charts and chemical structures do not exist at all as far as traditional tooling is concerned. A researcher can be one floor away from the answer to their question and have no realistic way to find it.

Every large organisation has some version of this. It rarely appears on a risk register, because nothing is broken. The work simply gets redone, quietly, at a cost nobody totals up.

What they did — three stages, in order

What makes this case instructive is what they did not do. They did not leap to a flashy AI product. There was no assistant launched in quarter one to demonstrate momentum. They followed a deliberate path, and the order is the substance of it.

Stage one — build a foundation a model can read

Restructure the unstructured, siloed, scanned material into a unified enterprise knowledge base that a large model can actually read. No intelligence yet. No demo. Just the unglamorous work of making the asset machine-readable.

This stage solved search efficiency on its own, which is worth noticing: it delivered real value before any AI was applied. That is what makes it politically survivable. A foundation programme that produces nothing for eighteen months does not survive its second budget cycle, no matter how correct it is.

Stage two — retrieval-augmented generation

A researcher asks a question in natural language; the system retrieves the relevant material and returns a structured answer. This is the stage everyone wants to start at, and it only works because stage one happened. Retrieval over a corpus that is half scanned images retrieves half the corpus — and, worse, gives no indication that it has done so.

Stage three — agents that act

Agents that draft compliance documents and handle routine text work. The shift from answering to doing. Note that this is the third stage, not the first, and that each preceding stage is load-bearing for it.

The part most write-ups skip: strict workflow underneath

Here is the design decision I find most instructive, and it runs against the current fashion.

The front end feels like a simple chat interface. Behind it runs a strict workflow orchestrated with LangGraph. Tool calls follow defined paths. Agent identity is tightly controlled. Behaviour stays stable, controllable and traceable — while the model’s flexibility is tapped only where semantic reasoning is genuinely needed.

This is the mature answer to a question most organisations are currently getting wrong in one of two directions. Either they forbid autonomy entirely and get a search box with better manners, or they hand an agent broad latitude and discover that in a regulated environment they cannot reconstruct why it did what it did.

The resolution is not a setting on a slider. It is architectural: constrain the path, free the reasoning. Use the model where ambiguity genuinely requires judgement, and use deterministic orchestration everywhere else. In a regulated industry, traceability is not a nice-to-have you trade against capability — it is the licence to deploy at all.

The real lesson

For most traditional enterprises, digitalization was never truly finished. It was declared finished. Programmes ran, systems were bought, a transformation office was wound down, and a great deal of the actual work — the archives, the edge cases, the departmental spreadsheets, the scans — was quietly left where it was, because the value of finishing it could not be articulated.

AI transformation is essentially forcing companies to finish that unfinished work. This is why so many programmes stall in a way that feels inexplicable to the people running them: the AI is fine, the use case is real, and the thing underneath was never built. Without a solid digital foundation, no impressive Q&A bot or agent workbench solves the underlying problem. It relocates it, and adds a licence fee.

There is something almost fair about it. The bill for a decade of deferred foundational work has arrived, and it arrived disguised as an innovation programme.

One question before anyone shows you a demo

Can your enterprise’s knowledge assets be read efficiently by a large model?

If the answer is no, you now know what your first stage is — and it is not the one on the roadmap. That is a difficult conversation, and it is considerably cheaper than the alternative, which is discovering the same thing two years and several vendors later.


📚 Start here: AI-Native Organization Design — the full research hub, with every article and video in one place.

More in this series


Watch the full series

This article accompanies a video from AI-Native Operating Model and Org Design — a series on how organisations actually absorb AI, and where they break.


Chunfeng “Breeze” Dong is an executive coach (ICF PCC, CPCC) and founder of Springbreeze Ventures, with twenty years in organisational development inside Fortune 100 companies — Roland Berger, Siemens, ABB and Roche. She writes on AI-native organisation design, human–agent governance and change leadership.

📘 The Living Organization · 📘 A Soulful Transition · 🔗 LinkedIn