AI is exposing manufacturing’s information governance debt

Manufacturers are moving quickly to give employees access to copilots, AI assistants and enterprise automation tools, but the information those systems are being asked to work with was rarely created with AI in mind. Engineering documents sit across shared drives and collaboration platforms, permissions have accumulated over years, multiple versions of the same file remain accessible, and business context is often understood by people rather than encoded in the information itself.

M-Files, an information management platform that helps organizations organize, govern and access business content, argues that AI is exposing weaknesses that were already present rather than creating an entirely new class of problem. Its recently launched AI Readiness Model assesses organizations across strategy, metadata, governance, process automation and AI activation, with particular attention on whether the information underneath AI systems is accurate, controlled and understandable enough to support trusted decisions.

For Tiina Hietajarvi, Principal Manufacturing Industry Adviser at M-Files, one of the risks is that manufacturers can move too quickly from experimentation into business use without defining what AI should be allowed to access, how employees should use it or whether the underlying model is appropriate for the task.

“What I see many manufacturing companies doing is distributing AI or Copilot licenses to personnel without necessarily training people on what they can do with them,” she says. “From a compliance perspective, the business may not even fully understand what it has agreed to with the vendor. Where is the data going? Is it being used for training? Manufacturers have always worried about vendor lock-in, and now there is another question around putting part of a business process into somebody else’s AI system when pricing, licensing and capabilities can change.”

The issue becomes more serious when AI is expected to provide precise operational instructions. Large language models are designed to generate responses rather than reproduce the same deterministic answer every time, which creates a different risk profile from conventional document retrieval.

“If you are using Copilot to read maintenance instructions from a PDF or Word document, there are two misconceptions,” Hietajarvi explains. “The first is that an LLM is not built to give the same answer repeatedly. The second is that it can be expensive to keep sending the model back to retrieve and read the same content every time somebody asks the same question. Manufacturers need to think about whether the task really requires generative AI or whether the answer should be controlled as part of the business process.”

AI amplifies weaknesses that people could previously work around

Traditional information environments often survived because employees knew where to look, which version to trust and who to ask when something was unclear. AI removes much of that human interpretation. Once an assistant starts searching across thousands of documents, old folder structures, inherited permissions and inconsistent naming become active inputs into automated answers.

The problem is especially visible in engineer-to-order environments, where documents move between engineering teams, subcontractors, suppliers and later into maintenance and service. PLM systems may contain part of the picture, while transmittals, quality records, installation information and supporting documents live elsewhere.

“In an engineer-to-order process, you may have thousands of engineering documents that need to go to the correct vendor in the correct version and with the right compliance chain,” Hietajarvi says. “Historically, a lot of that has been handled through network drives, email or even USB sticks. The context is what allows you to know which project the document belongs to, which subcontractor is involved, who the engineer is, which component or product family it relates to and how that information eventually connects into the installed base and maintenance.”

M-Files describes one consequence of weak controls as the “Copilot liability zone”: information may have been overshared for years through outdated folder permissions without causing an obvious incident, but AI can suddenly surface it to users who were never intended to see it.

Manufacturing creates particularly sensitive examples because documents can contain intellectual property, pricing, contractual terms, quality information and evidence of unresolved production issues. “You would not want engineering IP distributed to people who were never meant to see it,” Hietajarvi adds. “Vendor contracts are another obvious compliance risk because AI might suddenly be able to read pricing or commercial terms that were previously restricted. Quality information can also be sensitive. If there is an ongoing quality issue with a supplier or within a factory, somebody without the right context could see that information, misunderstand what it means and not know what corrective actions are already under way.”

Folders create another difficulty because location does not necessarily establish status. An AI system may find several versions of a file without understanding which has been approved, which is obsolete and which is still being worked on. Hietajarvi argues that information needs business context around it, including approval status, project, product, owner and relevant lifecycle stage, rather than relying on the model to infer all of that by reading the document itself.

Trusted AI needs provenance as well as permission

Security is only one part of the trust problem. Manufacturing users also need to understand why an AI system produced a particular answer, especially when the output affects maintenance, quality, engineering or another process where small inaccuracies have physical consequences.

Hietajarvi believes trust can break down quickly when experienced employees encounter an answer that appears plausible but cannot be traced back to its source. A maintenance technician who notices that an AI-generated instruction has reordered two steps may stop using the tool altogether, even if most of its responses are correct.

“If Copilot checks a PDF maintenance instruction and gives you steps one, two, three, four, seven, five, eight, an experienced person will immediately say the system is unreliable,” she explains. “If it is just a query to an LLM, it is effectively a black box. No expert is going to trust a black box for something that needs to be exact. It can be useful for conversation and ideation where you are going to check the result, but that is very different from a controlled business process.”

A stronger approach is to make the source and reasoning visible. Repeated process instructions can be treated as governed business rules rather than regenerated from a document every time, while an AI agent can show which approved information, product record or customer context it used to reach an answer.

“You need to be able to explain where the result came from and why,” Hietajarvi says. “If the system can show that it looked at the approved information for this product, this customer and this process, and that the information passed through a human approval step before it entered the system, you have a much better chance of overcoming that trust barrier. The challenge is not only technical; it is getting people to trust and use the result.”

Data quality remains the underlying weakness even in organizations that believe they have already addressed information management. Metadata may exist but be incomplete, approval chains may not be enforced consistently, and operational documents may still sit outside governed systems. AI simply makes those inconsistencies more visible because the quality of the response depends on the quality of the context available to it.

Governance cannot become another information silo

Better control does not mean restricting every document to the narrowest possible audience. Hietajarvi argues that manufacturers may need to rethink long-established departmental boundaries if they want AI to create value across engineering, production, maintenance and research.

“Manufacturing companies are ultimately human systems, and some restrictions are more about organizational politics than genuine risk,” she says. “If you have proper control, does it really matter if research can see relevant maintenance records or engineering information? The question should be whether the information can be shared safely with the people who have a legitimate reason to use it, rather than simply keeping every department inside its own data boundary.”

That requires more granular governance than giving somebody access to an entire folder. Permissions can be shaped around product lines, countries, projects, roles or specific business objects, allowing AI to cross departmental boundaries without making sensitive information universally visible.

The longer-term opportunity is broader than document search. Hietajarvi sees contextual information connecting engineering, ERP, PLM, CRM, project milestones, quality records and maintenance history so that AI can understand relationships across the lifecycle rather than treating each document as an isolated file.

“With the right context, AI can start learning across the lifecycle,” Hietajarvi concliudes. “You can connect engineering documents, project milestones, quality information and supplier delays, then look historically at similar product families or projects. That creates the possibility of predicting maintenance issues or revenue risk from delays because the AI can follow the connected chain of information. Without that chain, it cannot really do that.”

Manufacturers therefore face a less glamorous AI challenge than selecting the latest model or assistant. Before copilots and agents can be trusted with increasingly important work, organizations need to know what information they hold, which version is authoritative, who should be able to see it and what business context gives it meaning.

AI can make information easier to find and decisions faster to reach, but speed magnifies whatever already exists underneath. In a poorly governed environment, that means faster access to inconsistent, outdated or overshared information. In a well-governed one, the same technology can become a route to more trusted decisions across the manufacturing lifecycle.

!-- Impression Tag --> Ad