AI makes statistical thinking more important

!-- Impression Tag --> Ad

A production process can drift upward on a chart without anything being wrong. Feed the same movement into an AI system and it may produce a confident explanation and recommendation to intervene. The danger is not that the technology has failed to recognize a pattern, but that nobody has established whether the pattern matters.

Manufacturing has spent decades learning how to distinguish normal variation from meaningful change. Methods such as statistical process control and regression exist because operational data is noisy and a plausible explanation is not the same as a reliable conclusion. AI does not make that discipline obsolete. By putting sophisticated analysis in front of far more people, it makes statistical judgment more important.

Generative AI can summarize a dataset in seconds and describe the result in authoritative language. Fluency cannot establish causation or prove that an apparent relationship will survive changing production conditions. Faster answers are useful only when the evidence behind them remains visible.

Explanation is not prediction

Part of the confusion comes from the way AI now covers very different technologies. Statistical machine learning and large language models are often discussed as though they perform the same function, even though one may model process behavior while another helps someone understand the result. Treating them as interchangeable encourages manufacturers to mistake a better explanation for a better prediction.

For Minitab President and CFO Josh Zable, the current excitement around generative AI has blurred that distinction. Large language models can explain what a user is seeing, but statistical methods remain the foundation for prediction and process improvement. “You really have to think about AI in two ways,” he says. “One is statistical machine learning methods, and one is what we now call large language models. Large language models are great at explaining what you see or what is happening, but they are not great at predicting what will happen. Statistical analysis and machine learning help you improve processes, reduce variation and predict outcomes. Large language models help you interact with that information. If you have not implemented statistical methods or machine learning methods and you jump straight to the large language model, you are missing a lot of the value creation.”

Most factory questions are not requests for a description. A quality engineer needs to know whether a process has moved out of control, while a maintenance team must distinguish developing damage from routine variation. An AI-generated narrative may help interpret a control chart, but it cannot decide whether the right data was collected.

The signal may not be in the data

Industrial AI is often presented as a way to extract value from the huge volumes of information generated by connected equipment. Volume does not compensate for measurements that are incomplete, poorly sampled or incapable of capturing the phenomenon being investigated. An algorithm cannot recover a physical signal that a sensor never recorded.

Drew Mackley, Director of Sales Enablement for Reliability Solutions at Emerson, sees this limitation regularly in asset monitoring. Wireless sensors and cloud analytics have widened access to condition data, but the underlying measurement still determines what the system can know. “If your sensor only goes to one thousand hertz and the gear mesh frequency at the speed it is turning is above one thousand hertz, you cannot even measure it,” he says. “No amount of AI is going to help you determine a gearbox failure there. Different faults can also be directional, so you may see something in the axial direction but not in the vertical. There is a great deal involved in machinery health and vibration analysis that standard pattern recognition may not yet have fully developed. AI is a great starting point and it will create efficiency, but we still need to double-check the result and get a second opinion because there is a lot to consider.”

Models do not automatically know whether the sampling interval, sensor position or operating state was appropriate for the question. Continuous monitoring helps, but readings still need to be considered alongside process conditions. A vibration increase caused by cavitation may originate in a valve restricting flow rather than in the pump itself, so a credible pattern can still lead to the wrong diagnosis. Statistical thinking forces teams to ask what was measured and with what degree of confidence, preventing manufacturers from acting decisively on the wrong signal.

Prediction still needs a test

Even a statistically credible prediction does not show what will happen after an intervention changes the system. Factories contain interdependencies and constraints that can produce unintended consequences once a decision moves from analysis into operations. Speeding up one production stage may only create a larger bottleneck elsewhere.

Simulation offers a way to test those consequences before production absorbs the risk. Historical data provides the starting point, while distributions and operational knowledge allow the model to represent uncertainty rather than assuming the next shift will resemble the last. The recommendation becomes a hypothesis to examine rather than an instruction to follow.

Tony Smith, Solutions Engineer at Simul8, argues that simulation matters because manufacturing variables do not remain fixed. “Order schedules will change, deliveries from suppliers will change and the quality of the materials arriving from those suppliers will change, so the variables move from day to day,” he explains. “You need to see what those changes could do before production goes live. The purpose of simulation is not to analyze what happened in the past. It is to predict what is likely to happen under different conditions in the future, so people can make decisions earlier and with greater confidence.”

A revised schedule may look optimal until labor availability or cleaning cycles are introduced. Simulation makes assumptions visible because teams must agree on how the process behaves and where uncertainty sits. It can expose when an improvement in one area creates pressure elsewhere.

Statistical literacy becomes an operating skill

Natural-language interfaces will let more employees question industrial data without waiting for an analyst. That matters particularly for smaller plants without large data science teams. It also means people with limited statistical training will be asked to judge increasingly sophisticated outputs.

Zable uses the control chart as a simple example. Many operators can recognize a line moving upward without understanding whether it represents normal process variation or a statistically significant change. “If you have a company using a control chart to understand whether its process is in control, that is a simple use case that has been around for decades,” he says. “Many people do not really understand that it relates to a process with normal variation and that you are trying to identify the abnormal variation. A large language model can help an operator analyze that control chart and explain the result, but you still must know enough to collect the data and set up the control chart correctly in the first place.”

Manufacturers do not need to turn every supervisor into a statistician, but they do need enough statistical literacy to challenge an output that conflicts with process knowledge. Domain expertise remains central because every process contains physical relationships that a generic model will not fully understand. Engineers and operators must still decide whether a pattern reflects a real mechanism or an accidental association.

From plausible answers to defensible decisions

The strongest industrial analytics projects are judged by what changes in the operation, not by how convincing the output sounds. Predictive insight should lead to fewer failures or a measurable improvement in the use of energy and materials. It needs to survive contact with the physical process.

At TC Energy, predictive analytics was combined with contextualized operational information to improve fuel efficiency and reduce carbon dioxide emissions by 54,000 metric tons, according to Mounir Boemond, Global Director of Sustainability Value at AVEVA. The result did not come from asking AI to interpret an isolated dataset, but from bringing operational information together in a form that could support an engineering decision. “The advantage of machine learning and AI is that it helps speed things up, but the data must be contextualized and linked to the operation,” he notes. “If you are not integrating the data and giving it meaning, then you are not getting the information needed to make the right decisions. The technology can accelerate the process, but it still must be grounded in what is happening on the plant floor.”

This is the standard manufacturers should apply as AI becomes embedded in analytics platforms. The issue is not whether the system can produce an answer quickly or express it clearly. The method must distinguish signal from noise and remain valid when conditions change, while allowing the proposed action to be tested before the factory absorbs the consequence.

AI will make statistical analysis easier to use and more widely available, but it may also encourage people to accept conclusions they would once have challenged. Manufacturers that retain their statistical discipline will gain speed without surrendering rigor. The factory of the future will not be run by intuition alone, but neither should it be governed by whichever model produces the most confident explanation. Competitive advantage will belong to manufacturers that can move from plausible answers to defensible decisions, and that still requires people who understand what the numbers are saying.

!-- Impression Tag --> Ad