.png)
In 2025, researchers published a study of a problem I think every network team supporting automation will eventually face. The study, co-authored with a researcher from Hyundai Motor Group's Manufacturing SW Platform R&D Team, analyzed AGVs operating in a large automotive factory — dozens of access points, several dozen vehicles, 3.5 months of production data, over 1,100 disconnection events, 331 of them examined in depth [1].
The factory had monitoring logs showing roaming and disconnection events. But the researchers noted that conventional network-level indicators such as signal strength, noise, and throughput cannot reveal the underlying cause of a disconnection. They show that performance has degraded, but not whether the failure originated in roaming, authentication, the AP, the AGV's Wi-Fi chipset, its firmware, or its driver. To identify the direct cause of each disconnection pattern, the researchers had to pull kernel logs from the vehicles themselves and read them against the network evidence. One finding stayed with me: even a technically successful roam did not always produce a stable operational connection.
I have heard the same pattern from teams managing connectivity for automation applications. The network had data. The operation still had failures. What was missing was the reasoning chain between them.
The question is no longer only whether the network is healthy. It is whether the operation depending on it is working.
In operational environments, "up" is often only a component-level description. The AP may be reachable, the controller may report healthy status, and the device may appear associated while the end-to-end operational service is already broken.
Network operations teams are built to answer important questions:
But operational systems measure a different set of outcomes:
These answers may already exist in the fleet manager, WMS, inference platform, or MES. What is usually missing is the connection between those outcomes and the infrastructure supporting them.
Network metrics remain essential. But in environments where physical operations depend on connectivity, they are only one layer of the answer.
Network availability is a technical condition. Operational continuity is the outcome.
This is the shift from network assurance to operational assurance: connecting network behavior to the device, application, and physical workflow — and being able to explain where and why that chain broke, so the same failure does not keep recurring.
I think about this as four connected layers.
Infrastructure: Access points, radios, switches, controllers, private 5G RAN and core, DAS, gateways, and backhaul. This is where most network monitoring is concentrated.
Device: Whether the endpoint can connect, authenticate, roam, maintain its session, and recover. Firmware, drivers, antennas, and device-specific behavior all matter here — and as the factory study showed, this layer can expose causes that infrastructure telemetry alone cannot see.
Application: Whether the required transaction or service completed. Did the AGV receive and acknowledge its command? Did the video reach the inference engine? Did the scanner transaction reach the warehouse system?
Operation: Whether the physical workflow succeeded. Did the robot finish its route? Did production continue? Were packages processed at the required rate?
The dependency chain is simple:
Infrastructure → Device → Application → Operation
The model is not strictly linear. Security, identity, timing, and physical context cut across all four layers. It is a way to organize the evidence, not a strict causal sequence.
But the operating environment rarely reflects even that simplicity. Each layer often belongs to a different team, vendor, and tool. The network team sees RF and infrastructure. The device vendor sees the endpoint. The application team sees timeouts. Operations sees robots stopping.
Everyone may be correct within their own boundary while no one can explain why the workflow is failing.
In many automated environments, the operational outcome is already recorded. The fleet manager, WMS, MES, vision system, or location platform may already hold it. The operational outcome is not missing. It is disconnected from the evidence needed to explain it.
The harder problem is connecting that outcome to the device, physical location, network state, validated baseline, and changes that preceded it. Each system has identifiers within its own domain, but no single domain standard provides a common end-to-end correlation model across the operational workflow and the network supporting it. At many sites, connecting them still requires custom integration using asset identity, location, timestamps, order or job identifiers, and site-specific naming conventions. This is why "green dashboards, red operation" persists even when every data source is theoretically available.
This is where a site knowledge graph earns its place. It is not a place to copy every raw event. It provides the identity and relationship layer that maps those records to the same assets, locations, workflows, and incident timeline — so that when the workflow fails, the question of what changed can be answered across silos rather than within one.
Suppose the original design assumed an open zone, a walk test validated it six months ago, and a new mezzanine was later installed. The graph should make that physical change visible alongside the timing of new failures, the firmware version running on the affected vehicles, and the access points serving the area. The underlying data still lives across design, survey, network, device, application, and operational systems. The graph does not replace them. It connects their entities.
It also has to be honest about its own freshness. A knowledge graph can go stale just like an as-built drawing. The difference is that a good one knows it — it can tell you the RF baseline is nine months old, or that the current floor plan conflicts with the last survey.
The operational-assurance stack provides the layers. The knowledge graph provides the cross-domain connections that no single domain standard provides.
Now consider an agent given a clear objective: investigate why AGV 41 failed to complete Route 7.
An agent connected only to a WLAN controller can inspect associations, signal quality, roaming events, retries, and configuration. That may be useful, but it sees only one part of the system.
An agent grounded in the site knowledge graph can ask a broader sequence of questions. Where was the vehicle when it failed? Which infrastructure was serving it? What performance was predicted in that area? What did the most recent survey validate? Was the device firmware changed? Did the authentication policy change? Was the application session lost before or after the roam? Have similar failures occurred along the same route?
Instead of searching every available log, the graph narrows the reasoning space to the dependencies surrounding the failed operation. It does not prove causality, but it makes the relevant relationships explicit and identifies which evidence is still missing. The agent can form competing hypotheses, compare the current state with the validated baseline, and recommend the next test. It can also produce something that is often missing in complex incidents: a defensible explanation of whether the likely cause sits in the network, device, application, security policy, or physical environment.
This is not a real-time control loop. Deterministic control and safety systems close to the process still handle real-time safety. The agent's job is explanation, attribution, and pattern detection across incidents — work that today happens days late, or does not happen at all.
AI can shorten the path from symptom to cause, but only when it can see the dependency chain.
A warehouse-automation case documented by Belden makes the commercial impact clear. The operator used AGVs and autonomous mobile robots to process packages. When safety packets were missed over the wireless link, the robots were designed to enter a safe state. That preserved safety, but it also created downtime. The company could not meet its package-per-hour processing rates, and its customer uptime commitments carried financial penalties [2].
The dependency chain was: missed communication → robot enters safe state → throughput falls → customer commitment is missed.
The customer was not ultimately buying access-point uptime. It was buying packages processed safely per hour.
The network SLA was only one layer of the real SLA.
Since Edition 3, my conversations with SI leaders keep circling the same question: what should a managed-services offering actually deliver in this era? I see a progression.
Network assurance focuses on infrastructure health, RF performance, availability, and configuration.
Experience assurance adds device behavior, roaming, authentication, session continuity, and application transactions.
Operational assurance connects those layers to the physical workflow, validated baseline, site changes, and business outcome.
This does not mean the SI becomes liable for every robot, application, or production failure. It means the SI can provide defensible fault attribution. The answer may be that the network is the cause. It may be that the network is operating within its validated envelope and the issue is in firmware, application timing, security policy, or a physical change at the site. The value is the ability to show the evidence and identify where the dependency chain most likely broke.
That moves the managed service from infrastructure-status reporting to operational assurance.
The researchers in the automotive factory did not explain the failures by looking harder at signal strength alone. They needed evidence from the network, the vehicle, and the interaction between them.
Physical AI will make this pattern more common. Robots, cameras, scanners, medical devices, and industrial systems do not experience the network as a collection of infrastructure KPIs. They experience whether their task completes.
Network health remains essential. But it is no longer the final measure. The network can appear healthy while the operation is failing — and the next operating model has to see the infrastructure, device, application, and physical workflow as one connected system.
A green network is not the goal. A working operation is.
[1] Bae, S., Shin, J.H., Paek, J., "Diagnosing AGV Wi-Fi Disconnection via dmesg Logs: A Real-World Factory Case Study," ICTC 2025. https://nsl.cau.ac.kr/papers/ictc/ictc2025-bae-wifi-dmesg.pdf
[2] Belden, "Eliminating Downtime and Timeouts in an Automation Company's Fully Autonomous Warehouse," 2024. https://www.belden.com/knowledge-hub/resources/case-studies/eliminating-downtime-and-timeouts-in-an-automation-companys-fully-autonomous-warehouse