AI is exposing the factory’s integration debt
Every point-to-point connection on a factory floor probably made sense when somebody installed it. A maintenance application needed machine hours, a historian needed process data, ERP needed production information, and another project needed a few tags from a PLC. Repeat those reasonable decisions over ten or 20 years, across plants with different equipment and local engineering practices, and the result is an architecture nobody would deliberately design today.
Artificial intelligence is making the consequences harder to ignore. Aron Semle, CTO of HighByte, an industrial software company focused on Industrial DataOps, says manufacturers are discovering that AI cannot simply be connected to decades of fragmented operational data and expected to understand it.
“If you look across ten, 15 or 20 years, you find little pieces of connectivity between all these different systems,” he says. “Then you go to the next plant, and it has a completely different OT stack, or perhaps the same stack implemented differently. You end up with this spaghetti architecture across many factories. Now executives are asking what they are doing with AI, and the technical teams are realizing: what are we going to connect the AI to? We do not have the data foundation to do this.”
AI changes the scale of the problem because useful industrial decisions rarely depend on machine data alone. A process value may describe what an asset is doing, but understanding why performance has changed could require maintenance history, production context or laboratory information held somewhere else. Semle argues that manufacturers therefore need to move beyond connectivity as the objective and make existing information understandable and reusable without replacing the systems already running production.
The missing layer is context
Raw industrial data carries far less meaning than manufacturers sometimes assume. A value such as 1,025 tells an AI model almost nothing until it knows what the number represents, which machine produced it, where that asset sits in the process and what was being manufactured at the time.
“I sometimes put a number on a slide and ask the audience what it means,” Semle says. “There is no answer. That is effectively the context an LLM has if you just give it raw industrial data. Then you start building it out: it is a cycle count on an injection molding machine on line three. What raw material is it running? What is being produced? What happens when those parts are tested downstream? You must build up from the raw data points to something that describes what is happening.”
“You cannot come into a plant and start ripping out PLC connectivity, HMI, SCADA or the systems that are running the business today,” Semle says. “You need to hook into what is already there and provide an augmentation layer. It connects to the existing systems, gives the data a common taxonomy, contextualizes and governs it, and then provides access. The consumer should not have to understand all the complexity that has been built underneath over the previous 10 or 15 years.”
That approach turns data architecture into a layer of abstraction rather than another replacement project. Existing operational systems continue doing the jobs they were installed to do, while newer applications gain a more consistent way to discover and use the information spread across them.
Unified namespace architectures have helped manufacturers confront some of the same fragmentation, but Semle cautions against confusing the MQTT broker at the center of many UNS implementations with the data architecture itself. Most of the difficult engineering happens before information is published.
“In practice, 80 to 90 percent of the work is connecting to the existing systems and contextualizing the data,” he says. “Publishing it to an MQTT broker is the last stretch and technologically that is relatively trivial. The good news is that the work you have already done building that contextualized information can now be reused for LLMs. You were building a data foundation whether you called it that or not.”
AI readiness starts before the model
Connecting an LLM and producing an impressive demonstration can be relatively quick. Preparing the factory information so the model receives the right data in a usable form is where Semle sees most of the effort, partly because every plant contains knowledge that exists nowhere except in the people who understand its equipment and processes.
HighByte is using LLMs to help users map existing industrial information into contextualized models. An engineer can provide material such as P&IDs, PLC logic, an OPC server address space and the required data model, allowing AI to propose mappings that a plant expert can review and correct.
“If you have one expert in the plant, the question is whether you can use AI to multiply what that person can do,” Semle says. “The model can browse the OPC server, help link the tags and build the contextualized data, and then the expert reviews it almost as they would review code. They can quickly say this looks right, this does not and iterate. That use of AI for DataOps is very real, and customers are doing it today.”
Brazilian flat-glass manufacturer VIVIX shows how that foundation can support more advanced AI applications. The company used HighByte Intelligence Hub to merge and model information from OPC servers, SQL systems and other industrial sources before publishing contextualized data into its AWS environment. It is now developing a Virtual Engineer intended to reduce response time for customer quality incidents from days or weeks to minutes and cut the time spent finding and interpreting information by 90 percent.
Applications such as this point towards a more agentic future, but Semle draws a clear line between AI that is given controlled access to defined industrial information and autonomous agents allowed to roam across production systems. Despite the rhetoric surrounding autonomous factories, none of HighByte’s customers, he says, is currently allowing fully autonomous agentic AI to operate freely across industrial data.
“What is happening is much more controlled,” he says. “Customers are building internal chatbot tools and creating specific tools that allow AI to query precisely the information required for that use case. Instead of exposing an entire historian, they might allow it to query the event frames for a particular piece of equipment over a particular period. You limit the scope of what the AI can access, get more precise answers and generally improve performance.”
Standardize enough to scale
A common industrial data architecture becomes particularly important when manufacturers try to reproduce a successful application across plants. Two sites may use different machines and controls, but an application can still be portable if both expose the information it requires in the same way.
“If I build an agent in one plant and it delivers ROI on an injection molding machine, I want to take it to the next plant and connect it to a completely different machine but expose the same things,” Semle says. “If the machine state, cycle count and controls are represented consistently, I can be much more confident that I can replicate that application across plants and get the same value.”
Semle does not believe manufacturers need to wait for industry-wide standards defining every asset before making progress, nor should they spend a year designing a perfect enterprise model before anyone uses the data. He has seen centralized initiatives reach their first implementation only to discover that the model omitted information the business needed.
“You have to build as you go,” he says. “Start with an actual use case, build the best data model you can for it and implement it end to end. Let the business user work with the data, get the feedback and iterate as quickly as possible. Model definitions are always going to change as you learn more and bring in new equipment.”
As AI consumption grows, there is also an economic reason to become more selective about what models can see and when inference is genuinely necessary. Semle says large industrial companies are already starting to examine their AI budgets for 2027. Dashboards, reports and deterministic applications will continue to have a role because not every operational question warrants an AI call.
The rush towards industrial AI has exposed a problem that predates AI itself. Manufacturers are not short of machine data or connections. They are short of a consistent way to turn decades of fragmented information into something people, applications and AI can understand and reuse. The factories best positioned for industrial AI may be those that resist adding another clever connection and finally make sense of the ones they already have.

