Photo by Yuri Krupenin on Unsplash. Source: https://unsplash.com/photos/a-man-standing-in-front-of-a-control-panel-iK0ODqrNlvE (Unsplash License).

Executive Summary

Booz Allen Hamilton ran eight autonomous scenarios against an operational technology lab built to resemble a general manufacturing plant. Two unnamed frontier models reached the objective in every one. An agent moved from a perimeter compromise to actions inside the industrial control network in a little over sixteen minutes, and another found and moved a collaborative robot arm in minutes.

The result is not really about writing exploits. It is about the last step in the OT attack chain. Operator screens, motor speeds and physical equipment are now reachable by an agent handed an address range and nothing else, and the detection budget is measured in minutes.

The testers gave the models no source code, no engineering drawings and no operational technology coaching. The lab held programmable logic controllers, human machine interfaces, engineering workstations, a SCADA platform, a variable-frequency drive, sensors and a collaborative robot arm, assembled across mixed vendors and imperfect network segmentation. Every one of the eight objectives was met, and at machine speed.

Diagram of the four stage autonomous attack chain used in the Booz Allen Hamilton operational technology lab, from mapping the plant to moving a collaborative robot arm, with a note that agents reached the objective in eight of eight scenarios and no source code or engineering documents were provided.
The four stages the agents worked through, and the numbers that came out of the lab.

The models did not need the manual

On the first SCADA attempt the agents aimed at the wrong version of the operator interface and failed. So they changed the plan. They read the active sessions, worked out that the control room ran a different client build, found editable Jython code inside the exported project, rebuilt the payload and used the administrative interface to take over the screens the operators were watching. Specialised knowledge of odd protocols and proprietary hardware is no longer a barrier to that work.

One gateway handed over the whole plant

The finding with the longest tail is smaller than the robot arm. The SCADA gateway kept live pre-authenticated connections to 14 OT devices across two zones, so compromising that single host produced working access to every controller behind it. The agents noticed, in their own words, that a tag write through those connections was a path to every controller, wearing the SCADA server’s face.

That is a configuration decision, not a model capability. Most of the industrial estate in this lab is the sort of estate you inherit rather than design. Devices that carry no authentication or encryption, default vendor credentials on a cobot, and a relay still broadcasting for a communications peer that was decommissioned months ago. The agent that found the orphan device did it without being asked, noticed the misconfiguration on a safety-critical target, and proposed answering the dead address to collect a control channel.

Sixteen minutes is a detection budget

Volume-based defences assume a scanner replaying identical payloads. An agent that fails, reads the error, and rebuilds does not look like a scanner. It looks like an engineer. The related problem is the one our own coverage keeps finding in software supply and runtime layers, where one prompt was enough to walk out of an agent sandbox with credentials, and where thousands of GPU fleets still publish their own telemetry endpoints. AI infrastructure is being stood up faster than the isolation boundaries around it.

The practical output of the report is a number to plan against. If sixteen minutes is the time from a perimeter event to a controller action, then a once-a-shift log review is not a control, whatever the standard says it is.

Three questions for your own operation. Which of your OT hosts hold standing pre-authenticated connections to the rest of the plant, and who owns that list? How long would it take your team to notice a session that adapts after a failed attempt? And if a vendor default credential is still live on a device inside the safety perimeter, which team is accountable for it this quarter?

Sources. Booz Allen Hamilton published the lab findings in October 2026, and declined to name the two models. The Register carried the detail, including the agents that had to wait for human approval before any action with physical impact. CISA publishes industrial control systems guidance for the baseline practices the report says are still missing.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.