AI cannot run on yesterday’s factory data

A factory can generate millions of data points every second and still leave an AI system effectively blind. The problem is not simply whether the information exists, but whether it arrives while it still matters, carries enough context to be understood and can trigger a response before the production moment has passed.

That changes the role of industrial data infrastructure. For years, manufacturers concentrated on extracting information from machines and moving it into historians, dashboards and data lakes. AI and machine learning increasingly need a continuously updated view of the factory, plus a route back into operations when a recommendation requires action.

Magnus McCune, CTO of HiveMQ, sees the shift most clearly among manufacturers that have already spent years connecting equipment. “It is not sufficient to just move this data and land it in a data lake,” he says. “They want to act and gain insights from it while that data is in flight. Rather than realizing a day later that my batch failed, I want to understand in real time that it is happening so I can change the outcome.”

Data loses value when the factory has moved on

The traditional industrial data flow was largely extractive. OT systems produced information, IT pulled it upward, and analysis happened somewhere else. That worked when the output was a daily report or retrospective improvement exercise. It is less effective when an AI model is expected to influence a live process.

Latency is only part of the problem. Moving information away from the plant can strip away the knowledge required to interpret it. A temperature reading has little value if the system does not know which product was being made, which machine produced it or what operating state the equipment was in.

McCune argues that this is where centralized architectures often disappoint. “The person who understands that data does not usually work in that centralized system,” he continues. “They are the OT operator or shop-floor specialist at the edge. When I centralize the data, I can lose a lot of the insight about what that data even means. Then there is the latency in the return path. I may have done analytics in the cloud, but how do I take action on the shop floor?”

For AI, that return path matters as much as the journey into the model. An anomaly detected after a batch is complete becomes an investigation. The same anomaly identified early enough to alter temperature, pressure or speed becomes operational control.

McCune gives the example of a manufacturer using a large kiln whose electrodes gradually wear out. Failure during production can destroy the batch and create a safety risk from hot glass. Detecting the conditions that precede burnout gives the operator or automated system an opportunity to intervene.

AI needs a shared language for the plant

Speed without meaning simply moves bad information faster. One manufacturer McCune recently visited had spent years increasing data collection across roughly 200 facilities, producing around 100 million data points. The problem was no longer acquiring signals. It was making them understandable enough to support decisions consistently across the enterprise.

That is one reason manufacturers are paying more attention to unified namespace architectures. The important idea is not centralizing everything but defining a common structure and vocabulary so systems know what a piece of industrial data represents wherever it originates.

“For me, the unified namespace is a design pattern that helps create a single source of truth for what the data is,” McCune says. “Every site, every system and every AI agent cannot keep naming and integrating data in its own structure. Site one cannot be completely different from site two. The value is governing that consistency and creating a shared language so that when I refer to the temperature on a motor, people and systems know how to find it and what it means.”

At machine level, standards such as Sparkplug can add structure to raw MQTT data so equipment and applications can interpret it consistently. Higher in the architecture, manufacturers can apply their own enterprise models. The objective is not identical software everywhere, but data that remains legible outside the system that created it.

Ford is taking that approach as it modernizes a global manufacturing network containing thousands of machines per plant. Its legacy architecture made real-time data difficult to scale across sites, limiting common analytics and AI. The company adopted HiveMQ as the backbone of a centralized IIoT monitoring architecture, creating a more consistent stream of operational data for predictive maintenance, quality, asset utilization and energy use.

Gopalakrishnan Rajaram, Solution Architect for Industrial Systems at Ford, describes the effect: “HiveMQ facilitates real-time communication, allowing for immediate fault detection, calculation of key performance indicators and instantaneous alert generation directly back to the plant floor. This centralized approach streamlines data flow and operational oversight.”

Brownfield factories cannot wait for a clean sheet

The attraction of a perfectly standardized architecture collides quickly with manufacturing reality. Plants contain decades-old PLCs, proprietary protocols and equipment that may remain productive for years. Making the data AI-ready cannot depend on replacing the physical asset base first.

McCune describes HiveMQ’s approach as “additive, never rip and replace.” Older equipment can remain while the communications layer around it is modernized, translating legacy protocols into a governed data structure that can participate in the wider architecture.

That matters where replacing validated equipment carries its own cost and risk. Eli Lilly identified a connectivity gap across laboratories and manufacturing facilities where stand-alone instruments such as pH meters and balances were not integrated with computer systems. The company created an Equipment Connectivity Platform using MQTT and Sparkplug B to provide a standardized interface between equipment, MES and laboratory execution systems while also transferring data to the cloud.

Several hundred instruments have been connected through the platform, with controls supporting data integrity and reducing manual handling. The architecture has since become part of Lilly’s global connectivity strategy. The AI lesson is not that every instrument needs a model, but that future analytics become easier to deploy when information already moves through a consistent path.

Reliability is equally important because data infrastructure increasingly sits inside production. Mercedes-Benz uses HiveMQ within its Vehicle Diagnostics System to coordinate commissioning and testing of electronic control units during vehicle manufacture. The system spans 24 factories and around 10,000 test devices, handling nearly 470 million messages each month. If it is unavailable for too long, the assembly line can be affected.

Marius Hertfelder of Mercedes-Benz summarized the requirement plainly: “We have been running the VDS using HiveMQ for four years and the HiveMQ broker has not gone down. It is rock solid, completely reliable. This is very important since we cannot stop the factory assembly line.”

The AI-ready factory will govern data before it moves

Pilot projects can tolerate informal controls. A single site can keep a PDF defining its data structure and rely on a small team to follow it. That breaks down when the same architecture is deployed across dozens or hundreds of facilities.

McCune sees manufacturers moving from those “soft controls” toward mechanisms that enforce how data is named, structured, authenticated and routed. “When you have one site, it is reasonably easy to tell everyone that the data should look like this,” he adds. “When you go to 100 sites, that is no longer sufficient. You need the tooling, security capabilities and authentication mechanisms to ensure the data conforms to the model you have applied.”

This is also changing the relationship between IT and OT. The objective is not to collapse the two disciplines into one. OT still owns the constraints of the physical process, while IT brings experience in scaling, governance and enterprise integration. Central teams can define what trusted data should look like while plant teams retain responsibility for producing it from local systems.

AI makes that division more urgent because poor data can now influence decisions automatically. McCune recounts a packaging manufacturer that believed its sites were operating at around 80 per cent capacity. Deeper investigation suggested many were closer to 60 per cent. An AI system reasoning over the original figures could encourage unnecessary capital investment in new capacity.

The next manufacturing advantage will not come from collecting the greatest amount of data. It will come from creating a live operational picture that machines, people and AI can understand and keeping it close enough to production for decisions to matter.

“The trusted view wins every time,” McCune says. “The volume of data collected stopped being a reasonable differentiator years ago. It is much more about how you create understanding around that data, how you create trust around it and, specifically, how you act on it.”

AI may be forcing manufacturers to confront that architecture now, but the result is broader than AI readiness. Once factory data becomes contextualized, governed and continuously available from edge to enterprise, it stops being something the plant periodically reports. It becomes part of how the plant operates.

!-- Impression Tag --> Ad