RSC Topic: Risk Management & Risk Register

  • How often should we perform ISO 27001 risk assessments in a factory?

    ISO 27001 does not mandate a specific frequency for risk assessments, even in factory environments. It requires that risk assessments are defined, performed, maintained, and kept up to date. In practice, most industrial organizations combine a regular cycle (often annual) with event-driven reviews when relevant changes occur.

    Typical baseline frequency in factories

    In a regulated, brownfield manufacturing environment, a realistic pattern is:

    • At least once per year for a formal, documented risk assessment covering the scope of the ISMS (including key OT systems, if in scope).
    • Every 2–3 years for a deeper, structural refresh of the risk assessment methodology, asset inventory, and risk criteria, aligned with business and regulatory changes.
    • Event-driven updates whenever there is a material change or trigger that could affect risk.

    The precise cadence should be defined in your ISMS procedures and approved through your normal governance and change control processes.

    Event-driven triggers in a factory context

    Beyond the baseline cycle, ISO 27001 expects you to maintain risk information so that it reflects reality. In a factory, that means updating all or part of your risk assessment when:

    • New production lines, cells, or facilities are introduced, especially where new OT networks, controllers, robots, or IIoT connectivity are added.
    • Major changes to OT/IT architecture, such as new MES, SCADA, historians, remote access solutions, or cloud integrations.
    • Significant process or product changes that affect information flows, recipes, NC programs, or controlled technical data.
    • Security incidents or near misses, especially those involving production systems, quality data, or regulated records.
    • Major vendor or infrastructure changes, such as network segmentation projects, identity and access management changes, or decommissioning legacy servers.
    • New regulatory or customer requirements that materially change confidentiality, integrity, or availability expectations.

    In many cases you do not need to redo the entire risk assessment. You can perform a scoped, change-driven update and re-evaluate affected risk scenarios and treatment plans.

    Factors that should drive your cadence

    The “right” frequency is highly dependent on your context. Key factors include:

    • Rate of change in the plant: Rapid deployment of automation, connectivity, and data integrations pushes toward more frequent reviews (for example, semi-annual).
    • Criticality of operations: High criticality (safety-related production, aerospace or defense work, life sciences, long product liability tails) justifies a tighter cadence and more conservative triggers.
    • OT security maturity: Plants early in OT security and network segmentation work often uncover new assets and dependencies; a more frequent cycle helps keep the risk picture accurate.
    • Integration with other risk processes: If you already run regular safety, quality, or business continuity risk reviews, aligning ISO 27001 reviews to those cycles may be more sustainable than adding a separate schedule.
    • Regulator and customer expectations: Some customers or regulators will expect to see at least annual risk assessment evidence and show-how-you-updated-it-after-changes rather than a static, three-year-old document.

    Coexistence with legacy and mixed-vendor systems

    In brownfield factories, full system replacement just to “clean up” cyber risk is rarely viable due to validation burden, downtime risk, and qualification constraints. Your risk assessment schedule should reflect that reality:

    • Map legacy assets explicitly (old PLCs, unsupported HMIs, custom integrations) and re-check their risks whenever network topology or remote access changes.
    • Accept that compensating controls (segmentation, monitoring, procedures) are long-lived and must be re-evaluated regularly rather than assuming fast replacement of weak components.
    • Integrate with existing processes such as MOC, equipment qualification, and CSV/validation so that risk reviews are automatically triggered when validated systems change.

    A practical approach is to anchor your ISO 27001 risk assessment updates to existing plant change control workflows. Any change that would trigger re-validation, re-qualification, or a major MOC should also trigger a targeted information security risk review.

    Pragmatic minimums and tradeoffs

    If you need a concrete starting point for a typical factory with mixed OT/IT in an ISO 27001 ISMS scope:

    • Define a formal, documented risk assessment at least annually, with clear scope and methodology.
    • Specify in your procedure that material changes, incidents, or audit findings trigger a scoped re-assessment of affected assets and scenarios.
    • Align with budget and staffing: Very frequent full-scope assessments without adequate resources often lead to superficial results, which is worse than a well-executed annual assessment plus meaningful interim updates.

    Ultimately, the acceptable frequency should be risk-justified, documented in your ISMS procedures, and consistently followed. It should also be supported by actual evidence of updates over time, not just a stated policy.

  • Security controls

    Security controls are specific measures, mechanisms, or activities that an organization designs and applies to address identified information security risks. In the context of ISO 27001 and an Information Security Management System (ISMS), security controls are selected and implemented based on a documented risk assessment and risk treatment plan.

    Security controls can be:

    • Administrative (organizational): policies, procedures, roles, responsibilities, and governance structures that direct how security is managed.
    • Technical: logical or technological mechanisms such as access controls, encryption, logging, and network segregation.
    • Physical: measures that protect facilities and physical assets, such as locks, badges, and surveillance.

    Each control is defined so that it can be implemented, operated, monitored, and reviewed. ISO 27001 and its related guidance documents (such as ISO 27002) provide structured catalogues of control objectives and example controls that organizations can use when designing their ISMS.

  • Residual Risk

    Residual risk commonly refers to the level of risk that remains after all reasonably practicable controls, safeguards, and mitigations have been identified, implemented, and verified. In other words, it is the risk that is still present even after a manufacturer has applied its risk-reduction measures.

    In industrial and manufacturing environments

    In regulated manufacturing and industrial operations, residual risk is typically evaluated as part of formal risk management or hazard analysis processes. It appears in activities such as:

    • Equipment and process safety assessments (for example, machinery hazards, chemical handling, or automated line risks)
    • Quality risk management for products and processes (for example, risk of nonconforming product reaching customers)
    • Information security and OT/IT cybersecurity risk assessments
    • Data integrity and compliance risk analysis for MES, ERP, LIMS, and other regulated systems

    Operationally, residual risk is often documented in risk registers, Failure Mode and Effects Analyses (FMEAs), hazard analyses, or cybersecurity risk assessments. Each identified risk typically has:

    • Inherent risk: estimated before any controls are applied
    • Controls and mitigations: preventive, detective, and corrective safeguards
    • Residual risk: re-estimated after considering the effectiveness of those controls

    Residual risk is usually compared against an organization’s defined risk acceptance criteria. If the residual risk is above the acceptable level, further controls or design changes may be considered, or the situation may be escalated for management decision and formal risk acceptance.

    What residual risk includes and excludes

    Residual risk includes:

    • Risks that cannot be eliminated without fundamentally changing the process, technology, or product
    • Risks that remain due to practical limits on cost, technology, or feasibility of controls
    • Risks introduced by controls themselves (for example, complexity or new failure modes)

    Residual risk does not mean:

    • That no further risk exists after controls are in place
    • That the situation is inherently “safe” or “compliant”
    • That risks are formally accepted, unless this is explicitly documented through a risk acceptance process

    Use in workflows and systems

    Within industrial and regulated environments, residual risk is often:

    • Recorded and tracked in electronic quality management systems (eQMS), MES, or risk registers
    • Linked to specific controls such as standard operating procedures, digital work instructions, alarms, interlocks, or access controls
    • Re-evaluated after process changes, deviations, CAPA actions, or system upgrades
    • Used as input when prioritizing improvements, maintenance, or cybersecurity hardening

    Common confusion

    Residual risk vs. inherent risk: Inherent risk is the level of risk that exists before any controls are applied. Residual risk is the remaining risk after accounting for existing or planned controls.

    Residual risk vs. acceptable risk: Residual risk is a measured or estimated state. Acceptable risk is a threshold or decision. Residual risk may be judged acceptable or not, depending on defined criteria and documented justification.

    Residual risk vs. residual hazard: Residual hazard typically refers to the remaining hazardous condition (for example, a moving part that cannot be fully guarded), while residual risk considers both the hazard and the likelihood and severity of harm.

  • Consequence

    Consequence commonly refers to the outcome, impact, or result of an event, action, failure, or deviation. In industrial and regulated manufacturing environments, it is a core concept in risk, safety, and quality management, where consequence helps describe how serious a given hazard, nonconformance, or system failure could be.

    Typical uses in manufacturing and operations

    In operational contexts, consequence is often used as a structured way to describe impact in areas such as:

    • Safety and health: harm or injury to personnel resulting from an incident, unsafe condition, or equipment failure.
    • Product quality: effect on product performance, compliance to specification, or patient/end-user safety when a defect occurs.
    • Regulatory and compliance: impact on regulatory status, inspections, or legal standing if requirements are not met.
    • Business and operations: financial loss, downtime, scrap, rework, or delivery delays caused by disruptions.
    • Environment: impact on emissions, waste, or environmental releases from process upsets.

    In many risk methods, consequence is combined with likelihood (or probability) to estimate an overall risk level. For example, a risk matrix or FMEA will often rate consequence on a defined scale (such as negligible, minor, major, critical) for consistent evaluation.

    Consequence in risk and quality methodologies

    Several structured methodologies use consequence as a defined factor:

    • Risk assessments and HAZOP-style studies: consequence describes the severity of a scenario if a hazard is realized.
    • FMEA (Failure Modes and Effects Analysis): consequence aligns with the severity of the effect of a failure mode on the system, user, or process.
    • Process deviation and CAPA investigations: consequence helps classify deviations, nonconformances, or complaints (for example, critical vs. major vs. minor) and can guide prioritization.
    • Business continuity and supply risk assessments: consequence describes service disruption, revenue impact, or impact on critical customers if a risk occurs.

    Operationally, consequence ratings may be captured in MES, QMS, EHS, or risk register tools, and then used to drive workflows such as escalation, approvals, or specific controls.

    What consequence is and is not

    • It is an impact measure: what happens if an event occurs.
    • It is not the probability, frequency, or likelihood that the event will occur.
    • It may be qualitative or quantitative, depending on the method and available data.
    • It is usually defined relative to a specific context, such as safety, quality, or business impact, using agreed rating criteria.

    Common confusion

    • Consequence vs. likelihood: Consequence addresses “how bad would it be if this happened?” Likelihood addresses “how often or how likely is it to happen?” Many risk frameworks treat these separately.
    • Consequence vs. severity: In many practical risk and quality tools, these are used almost interchangeably. “Severity” often refers to the level of consequence on a scale, while “consequence” describes the broader effect.
    • Consequence vs. root cause: Root cause explains why an event occurred. Consequence describes what resulted from that event.

    Link to operational decision making

    Defined consequence categories and criteria support consistent decision making in industrial operations. For example, a higher-consequence classification may trigger more stringent controls, faster investigation timelines, specific documentation requirements, or management notification. Clear definitions and documented scales help ensure that different teams assess consequence in a similar way across sites, systems, and processes.

  • Near-Miss

    Core meaning

    A **near-miss** is an unplanned event or condition that had the potential to cause harm, loss, or other adverse outcomes, but did not actually result in injury, damage, nonconforming product, or reportable incident.

    In industrial and manufacturing environments, this commonly refers to situations where:

    – A hazardous condition was present and almost led to a safety incident
    – A process deviation nearly produced nonconforming product
    – An equipment or system failure was narrowly avoided before causing downtime or quality impact

    Near-misses are treated as early warning signals that a hazard, weakness, or control gap exists in the system.

    Use in industrial and regulated environments

    In operations and manufacturing systems, near-miss reporting and analysis is typically integrated into:

    – **EHS and safety programs**: Near-misses related to worker safety, machine guarding, lockout/tagout, chemical handling, or ergonomics.
    – **Quality management systems (QMS)**: Near-misses related to out-of-spec parameters, incorrect materials, or documentation errors that were caught before product release.
    – **Maintenance and reliability workflows**: Near-misses indicating potential equipment failures, such as overheated components, atypical vibration, or control system faults that auto-recovered.
    – **OT/IT and MES environments**: Events such as incorrect recipe selection, mis-scanned material, or unauthorized parameter changes that were detected by system controls before affecting production.

    Near-misses are often logged in event management or deviation systems, reviewed in risk or safety meetings, and used as input for root cause analysis and corrective or preventive actions.

    Boundaries and what it is not

    A near-miss:

    – **Does include** events where no actual harm or nonconforming output occurred, but where credible potential existed.
    – **Does not require** physical injury, environmental release, or confirmed defective product.
    – **Does not include** routine process variation that remains within defined limits and poses no credible risk.
    – **Does not include** purely hypothetical scenarios with no triggering event (those are typically handled in risk assessments, not near-miss logs).

    Near-misses may still involve minor consequences such as short pauses, alarms, or temporary rework, as long as the primary adverse outcome (e.g., injury, major nonconformance, or significant loss) did not occur.

    Data and system handling

    In digital operations and manufacturing systems, near-misses may be:

    – Captured as **event records** in EHS, QMS, or incident management tools
    – Linked to **equipment, batches, work orders, or locations** in MES or ERP
    – Categorized by **risk type**, **root cause**, or **process area**
    – Analyzed as **leading indicators** in dashboards and operations intelligence tools

    Some organizations use standard fields such as severity potential, likelihood, and classification (safety, quality, environmental, cybersecurity, etc.) to support structured analysis.

    Common confusion and terminology

    Near-miss is sometimes confused with related terms:

    – **Incident**: An event where harm, damage, or nonconforming output actually occurred. A near-miss stops short of that outcome.
    – **Hazard**: A source of potential harm that may exist independent of any particular event. A near-miss involves an event or situation in which the hazard nearly produced an adverse outcome.
    – **Risk**: The combination of the probability and consequence of an event. Near-misses are real-world occurrences that inform risk assessment but are not themselves risk ratings.

    In some safety literature, the term **”near-hit”** is used instead of near-miss, but the operational meaning is the same.

    Role in continuous improvement

    Near-miss information is frequently used as input to:

    – Problem-solving methods (e.g., 5 Whys, fishbone diagrams) to understand underlying causes
    – Risk reviews in quality or safety committees
    – Changes to procedures, training, or control strategies

    Because near-misses occur more frequently than actual incidents, they are commonly treated as important signals when monitoring the effectiveness of controls in manufacturing and other industrial operations.

  • NIST SP 800-37

    NIST SP 800-37 is a U.S. National Institute of Standards and Technology (NIST) Special Publication that defines the Risk Management Framework (RMF) for information systems. It describes a structured, lifecycle-based process for managing cybersecurity and privacy risk to federal information systems and organizations.

    The publication is formally titled “Guide for Applying the Risk Management Framework to Federal Information Systems” (current revision numbers may change over time). It provides process steps, roles, and decision points for selecting, implementing, assessing, authorizing, and monitoring security and privacy controls, typically in alignment with control catalogs such as NIST SP 800-53.

    Key elements

    Within regulated and industrial environments, NIST SP 800-37 is commonly referenced as a process model for managing cyber and information security risk to both IT and OT systems. Core elements include:

    • System categorization: Determining the impact level of a system based on potential harm from loss of confidentiality, integrity, or availability.
    • Control selection: Choosing appropriate security and privacy controls (often from NIST SP 800-53) based on the categorization and risk tolerance.
    • Control implementation: Implementing the selected controls in the system and its environment of operation.
    • Control assessment: Evaluating whether controls are implemented correctly, operating as intended, and producing the desired outcome.
    • System authorization: A formal risk-based decision by an authorizing official on whether to operate the system.
    • Continuous monitoring: Ongoing oversight of security posture, changes, and control effectiveness over the system lifecycle.

    Use in industrial and OT environments

    In industrial operations, NIST SP 800-37 is often used as a reference framework when:

    • Extending federal-style RMF practices to manufacturing OT networks, MES, SCADA, and process control systems.
    • Structuring how security controls (for example, from NIST SP 800-53) are selected, assessed, and monitored for plant systems handling regulated or sensitive data.
    • Aligning cybersecurity risk management with existing validation, change control, and quality management processes.

    Relationship to NIST SP 800-53

    NIST SP 800-37 and NIST SP 800-53 are closely related but address different needs:

    • NIST SP 800-37: Defines the overall risk management and authorization process (the “how”).
    • NIST SP 800-53: Provides a catalog of security and privacy controls that can be selected and applied (the “what”).

    In practice, organizations apply the RMF steps from NIST SP 800-37 and use NIST SP 800-53 as a primary source for control requirements, tailoring them to their specific systems and risk profile.

    Common confusion

    • Not a control catalog: NIST SP 800-37 does not list detailed security controls; it defines the risk management process that relies on separate control catalogs, commonly NIST SP 800-53.
    • Not limited to IT-only: While originally oriented to federal information systems, the RMF concepts are often adapted for OT, industrial control systems, and manufacturing execution environments, but this adaptation is organization-specific.

    Link to reassessment and monitoring

    In the context of NIST SP 800-53 control reassessment, NIST SP 800-37 provides the overarching lifecycle and continuous monitoring concepts that guide how frequently organizations review and update their controls. Reassessment intervals are derived from risk, impact level, and system changes rather than a fixed schedule defined in NIST SP 800-37 itself.

  • risk treatment plan

    A risk treatment plan is a documented plan that describes how an organization will address identified risks. It typically records the chosen risk treatment option for each risk (such as reduce, transfer, avoid, or accept), the specific controls or actions to be implemented, responsibilities, resources, and target dates, as well as how progress and effectiveness will be monitored.

    Scope and use in industrial and regulated environments

    In industrial operations and manufacturing, a risk treatment plan commonly covers risks related to operational technology (OT), information technology (IT), product quality, safety, cybersecurity, data integrity, and regulatory compliance. It provides a structured link between risk assessment results and the practical measures implemented in plants, systems, and processes.

    Typical elements include:

    • A reference to the risk assessment where the risk was identified and evaluated
    • The selected treatment strategy for each risk (for example, implement a control, modify a process, or formally accept the risk)
    • Descriptions of technical, procedural, or organizational controls to be implemented
    • Accountable owners and supporting roles for each action
    • Planned implementation dates, dependencies, and change control references
    • Criteria and methods for verifying that the treatment is implemented and effective

    In regulated manufacturing, the risk treatment plan is often expected to align with information security, quality, and safety management frameworks. It may be used as a key piece of evidence during audits to demonstrate that identified risks are being managed in a controlled and traceable way.

    Relation to standards and Annex-based control sets

    Where organizations use structured control sets (for example, those organized in annexes or appendices of security or quality standards), the risk treatment plan commonly provides the rationale for:

    • Which controls are selected or tailored to treat specific risks
    • Which controls are not applied and why (for example, not relevant or addressed by alternative measures)
    • How selected controls are implemented across IT, OT, MES, ERP, and other manufacturing systems

    In brownfield or legacy environments, a risk treatment plan may explicitly capture coexistence with older systems, compensating controls, and formal change control steps required to avoid disrupting production.

    What a risk treatment plan is not

    • It is not the same as the risk assessment itself. The assessment identifies and evaluates risks; the risk treatment plan documents how they will be addressed.
    • It is not only a high-level policy statement. It should contain actionable, traceable activities, not just general intentions.
    • It is not limited to cybersecurity. It can cover quality, safety, supply chain, and other operational risks, depending on the organization's scope.

    Common confusion

    Risk treatment plan vs. Statement of Applicability: The Statement of Applicability typically lists applicable controls and their status. The risk treatment plan focuses on the concrete actions needed to implement or adjust those controls in response to specific risks.

    Risk treatment plan vs. risk register: A risk register records risks, their ratings, and sometimes owners. A risk treatment plan emphasizes the selected treatment options and implementation details. In some organizations these are combined, but the functions are conceptually distinct.

  • Critical Component

    Core meaning

    A **critical component** is a part, material, module, or subsystem whose failure, degradation, or incorrect performance can have a significant adverse impact on:

    – Safety of people, equipment, or environment
    – Product quality or product integrity
    – Regulatory or customer compliance
    – Continuity of manufacturing or business operations

    Critical components are typically identified through formal risk assessment, engineering analysis, or regulatory requirements, and they are subject to tighter controls than non‑critical components.

    Use in manufacturing and industrial operations

    In regulated and industrial environments, “critical component” commonly refers to items such as:

    – **Production equipment parts** whose failure can cause unsafe conditions or out‑of‑spec product (e.g., pressure relief devices, sensors on sterilization equipment, safety interlocks).
    – **Process control elements** that directly affect critical process parameters (e.g., temperature or pressure transmitters, control valves, PLC I/O modules on critical loops).
    – **Quality‑relevant components** of a product or assembly that directly influence key quality attributes (e.g., seals in sterile packaging, structural fasteners, critical dimensions).
    – **Compliance‑relevant items** required by regulation, customer specification, or standards (e.g., data integrity components in a GMP system, alarm systems for environmental monitoring).

    Once designated as critical, these components are often subject to:

    – Stricter change control and documentation
    – Enhanced supplier qualification and incoming inspection
    – Defined preventive maintenance and calibration schedules
    – Traceability and controlled storage or handling

    Boundaries and what it is not

    A critical component:

    – **Is not defined only by price or size**: low‑cost or small parts can be critical if their failure poses high risk.
    – **Is not necessarily safety‑critical only**: the impact can be on safety, quality, compliance, or availability.
    – **Is not a generic spare part category**: the designation is usually based on risk, not convenience or stock‑keeping.

    The term describes the **risk significance** of the component in its operational context, not its generic category in a catalog.

    Identification in typical workflows

    Industrial organizations commonly identify critical components through:

    – **Risk and reliability analyses** such as FMEA, HAZOP, or reliability‑centered maintenance studies.
    – **Process and product design reviews**, where engineers mark which parts are critical to function or quality.
    – **Regulatory or standards requirements**, which may mandate that certain items be treated as critical.

    The resulting critical component lists are then used by:

    – **Maintenance systems (EAM/CMMS)** to flag assets or parts with special maintenance and spare‑parts policies.
    – **MES and quality systems** to enforce traceability, inspections, and electronic records for designated parts.
    – **ERP and supply chain systems** to manage approved suppliers, lead times, and stock strategies for critical items.

    Common confusion and alternate usages

    “Critical component” is sometimes confused with related terms:

    – **Safety‑critical component**: a subset of critical components where the primary concern is human or environmental safety. All safety‑critical components are critical, but not all critical components are safety‑critical.
    – **Critical spare**: typically refers to spare parts kept in inventory because long lead time or unavailability could stop production. A critical spare is often (but not always) a spare for a critical component.
    – **Critical asset or critical equipment**: refers to whole machines or systems, whereas critical component refers to the parts within or used by those assets.

    Usage can vary between industries; some organizations formally define different classes such as quality‑critical, safety‑critical, and business‑critical components.

    Relation to risk and safety management

    Within risk and safety management frameworks, critical components are key objects for:

    – Risk control measures (engineering controls, alarms, interlocks)
    – Monitoring and inspection plans
    – Failure reporting and root cause analysis

    Documented lists of critical components help structure incident investigations and ensure that failures of these items are recorded, analyzed, and addressed through corrective and preventive actions.