Testing an office network and testing an operational environment are not the same exercise. An aggressive scan that is routine in IT can overload a fragile controller, disturb a serial gateway or trigger an unsafe state. Good OT security work begins by understanding the process, the authority structure and the boundaries of acceptable change.
Put process safety ahead of test coverage
Operations and engineering should help define the rules of engagement. Identify critical functions, safety instrumented systems, sensitive protocols, unsupported devices, maintenance windows, vendor dependencies and the conditions that require an immediate stop. The testing team should know who has authority to pause the work and how plant personnel will recognize test traffic.
The environment may include devices that cannot tolerate active probing, but that does not make assessment impossible. Architecture review, configuration analysis, passive traffic observation, account review and safe validation in a representative lab can often answer important questions before any production interaction.
Build a reliable asset and trust map
An inventory should show more than IP addresses. Record system function, owner, firmware or operating version where known, network zone, communications path, remote-access dependency, safety relevance and recovery constraint. Map how identity, vendor support, historians, engineering workstations and enterprise services cross the IT and OT boundary.
This map often reveals the most consequential weaknesses: a shared engineering credential, an unmanaged remote path, a flat trust zone, an obsolete protocol exposed beyond its intended segment or a backup that cannot be restored within the required operating window.
Use a graduated test plan
- Document review: diagrams, asset records, prior findings, change procedures and recovery plans.
- Passive observation: traffic and event review without introducing test packets into the process network.
- Configuration validation: exports, access rules and device settings reviewed away from live control paths where possible.
- Lab validation: controlled reproduction of suspected weaknesses against matching equipment or a faithful test environment.
- Approved production checks: narrow, reversible actions with operations present and an agreed stop condition.
Every stage should have a clear question. More packets do not necessarily create more assurance. A limited check that proves a dangerous trust path can be more useful than a broad scan whose operational effects are uncertain.
Report operational consequence, not just severity
Generic vulnerability scores are only one input. Decision-makers also need to know whether exploitation requires local access, whether compensating controls exist, what process could be affected, how easily a change can be made and whether remediation itself introduces downtime or safety risk.
A useful finding separates evidence from inference. It includes the observed condition, affected assets, plausible path, consequence, validation limit and a repair plan that operations can stage. When a condition was not actively exploited, say so. Do not present a theoretical path as a demonstrated compromise.
Evidence to retain
- Approved scope, rules of engagement and stop conditions.
- Asset and data-flow assumptions used to plan the test.
- Exact test source, timing, commands and tools.
- Operations log entries and alerts created by the exercise.
- Configuration snapshots supporting each finding.
- Exceptions, untested areas and reasons for exclusion.
NIST SP 800-82 Revision 3 describes OT security with explicit attention to performance, reliability and safety requirements. Its system-characterization and risk-management approach is a useful basis for scoping work that has to respect the physical process.
Authorization matters: Production testing should be authorized for the specific environment. Written permission, site safety rules, engineering supervision and recovery readiness are part of the technical method.