Root cause analysis seeks to explain why an incident occurred and why existing controls did not prevent, detect or limit it. The purpose is not to identify one person or one failed step. Most material incidents arise from several interacting conditions such as process design, technology, data, workload, governance, training and incentives.
A useful RCA distinguishes the immediate cause, contributing factors, control failure and systemic root cause. It tests the conclusion against evidence and similar incidents and produces actions that address the actual mechanism of failure. Naming 'human error' without analysing the environment rarely prevents recurrence.
This article provides a practical approach to RCA that can be scaled from routine events to complex cross- functional incidents.
Management question: Does the RCA explain why the event was possible and why the control environment failed, rather than merely describing what happened?
Why root cause analysis for risk incidents matters#
Weak RCA creates repetitive actions such as retraining, reminders or procedure updates while the underlying system or process weakness remains. Strong analysis supports control redesign, prioritises systemic action and allows management to identify common causes across incidents and issues. It also helps determine whether risk ratings, KRIs and assurance coverage need to change.
This topic is closely connected to Risk Incident Management Process: From Event Capture to Closure and Learning and Operational Loss and Near-Miss Management: Capture Better Data and Prevent Recurrence.
Core principles#
Establish facts before causes#
Build an evidence-based chronology and distinguish confirmed information from assumptions and missing data. The practical test is whether the organisation can apply this principle consistently when information is incomplete, ownership is distributed and decisions must be made within a defined governance timetable. In root cause analysis for risk incidents, a rule that exists only in a policy document is not enough. The rule should be translated into named data fields, accountable roles, review evidence and a clear exception path. Teams should be able to explain what was decided, who reviewed it, what information supported the conclusion and when the matter must be reconsidered. That discipline turns establish facts before causes from an administrative statement into an operating control.
Analyse multiple layers#
Consider immediate, contributing, control and systemic causes rather than stopping at the first plausible explanation. This element should be designed around the decision it is intended to support rather than around the convenience of a template. A sound approach defines the minimum information required, the acceptable source of that information, the person responsible for maintaining it and the reviewer who can challenge it. For incident investigators, operational risk teams, control owners and managers, the most useful outcome is not a larger volume of data; it is a reliable line of sight from the underlying risk condition to the management response. Where the condition changes, the record should show the new assessment, the reason for the change and any resulting action.
Use methods proportionately#
Apply five whys, fishbone, causal mapping, barrier analysis or change analysis according to complexity and severity. In practice, this requires both standardisation and room for judgement. Standardisation ensures that comparable risks are treated in comparable ways, while judgement allows context, materiality and emerging information to be considered. The balance is achieved through defined criteria, evidence expectations, approval thresholds and periodic review. Without those safeguards, root cause analysis for risk incidents can become either mechanically rigid or inconsistently subjective. A mature process makes the judgement visible and reviewable without pretending that every risk decision can be reduced to a single number.
Test the causal link#
Ask whether removing the proposed cause would probably have prevented or materially reduced the event. The design should also anticipate failure modes. Records may become stale, owners may change, thresholds may be interpreted differently and actions may remain open after their original rationale has expired. Controls therefore need due dates, reminders, escalation logic, independent review and closure evidence. For incident investigators, operational risk teams, control owners and managers, this is especially important because a weak follow-through process can create a false impression of control. The objective is to make unresolved exposure visible early enough for management to intervene.
Design actions against causes#
Corrective and preventive actions should address the causal pathway and include evidence of effectiveness. The practical test is whether the organisation can apply this principle consistently when information is incomplete, ownership is distributed and decisions must be made within a defined governance timetable. In root cause analysis for risk incidents, a rule that exists only in a policy document is not enough. The rule should be translated into named data fields, accountable roles, review evidence and a clear exception path. Teams should be able to explain what was decided, who reviewed it, what information supported the conclusion and when the matter must be reconsidered. That discipline turns design actions against causes from an administrative statement into an operating control.
A practical operating model#
1. Define the problem precisely#
Describe the event, scope, expected state, actual state, impact and time boundary. In practice, this requires both standardisation and room for judgement. Standardisation ensures that comparable risks are treated in comparable ways, while judgement allows context, materiality and emerging information to be considered. The balance is achieved through defined criteria, evidence expectations, approval thresholds and periodic review. Without those safeguards, root cause analysis for risk incidents can become either mechanically rigid or inconsistently subjective. A mature process makes the judgement visible and reviewable without pretending that every risk decision can be reduced to a single number.
2. Build the chronology#
Collect logs, documents, interviews, decisions, system events and control evidence in sequence. The design should also anticipate failure modes. Records may become stale, owners may change, thresholds may be interpreted differently and actions may remain open after their original rationale has expired. Controls therefore need due dates, reminders, escalation logic, independent review and closure evidence. For incident investigators, operational risk teams, control owners and managers, this is especially important because a weak follow-through process can create a false impression of control. The objective is to make unresolved exposure visible early enough for management to intervene.
3. Identify causal factors#
Analyse process, people, technology, data, governance, environment and external dependencies. The practical test is whether the organisation can apply this principle consistently when information is incomplete, ownership is distributed and decisions must be made within a defined governance timetable. In root cause analysis for risk incidents, a rule that exists only in a policy document is not enough. The rule should be translated into named data fields, accountable roles, review evidence and a clear exception path. Teams should be able to explain what was decided, who reviewed it, what information supported the conclusion and when the matter must be reconsidered. That discipline turns identify causal factors from an administrative statement into an operating control.
4. Validate root causes#
Challenge alternative explanations, compare similar incidents and review control design and operation. This element should be designed around the decision it is intended to support rather than around the convenience of a template. A sound approach defines the minimum information required, the acceptable source of that information, the person responsible for maintaining it and the reviewer who can challenge it. For incident investigators, operational risk teams, control owners and managers, the most useful outcome is not a larger volume of data; it is a reliable line of sight from the underlying risk condition to the management response. Where the condition changes, the record should show the new assessment, the reason for the change and any resulting action.
5. Agree action and learning#
Prioritise systemic corrective actions, assign owners and update risks, controls, KRIs and training. In practice, this requires both standardisation and room for judgement. Standardisation ensures that comparable risks are treated in comparable ways, while judgement allows context, materiality and emerging information to be considered. The balance is achieved through defined criteria, evidence expectations, approval thresholds and periodic review. Without those safeguards, root cause analysis for risk incidents can become either mechanically rigid or inconsistently subjective. A mature process makes the judgement visible and reviewable without pretending that every risk decision can be reduced to a single number.
Practical example#
A regulatory report is submitted with incorrect figures. The immediate cause is an erroneous manual adjustment. A five-whys review reveals that the analyst used a prior-period workbook because the current template was not clearly controlled. Further analysis shows that the reporting process depended on local files, the approval checklist did not confirm template version and the system change had not been incorporated into training. The root causes are therefore document-control and change-management weaknesses, not simply analyst error. Actions address template governance, automated validation and approval evidence.
The example is deliberately simple, but it illustrates an important point: a useful ERM process does not stop when a score has been produced. It connects the assessment to ownership, evidence, thresholds, actions, review and reporting. The resulting record should be capable of supporting management discussion without requiring the risk team to reconstruct the history from emails and spreadsheets.
Measures that show whether the process is working#
- RCA completion timeliness: Investigations completed within severity-based targets.
- Systemic-cause identification: Material incidents with causes beyond individual error or immediate failure.
- Action-cause alignment: Corrective actions explicitly mapped to validated causes.
- Repeat incident rate: Events recurring with the same or related causes after remediation.
- RCA rework: Analyses returned because evidence or causal reasoning was insufficient.
- Cross-incident themes: Common causes identified across products, units, systems or third parties.
Metrics should be interpreted together. A high completion rate can coexist with weak challenge, poor evidence or overdue remediation. Conversely, a temporary increase in identified issues may indicate that the organisation is becoming more transparent rather than less controlled. Management should therefore consider direction, materiality and the quality of response, not only the absolute number of exceptions.
Common implementation mistakes#
- Stopping at human error: The process, system and management conditions that made the error possible remain unchanged.
- Choosing a cause before evidence: Confirmation bias can shape the investigation and hide alternative pathways.
- Using one method mechanically: A simple five-whys chain may be inadequate for complex or multi-causal events.
- Writing causes as symptoms: Terms such as 'control failed' do not explain why it failed.
- Selecting easy actions: Training and reminders may be chosen because they are quick rather than effective.
These mistakes are avoidable when the operating model is designed before technology configuration begins. The organisation should agree terminology, ownership, approval thresholds, evidence expectations and reporting logic first. Technology can then enforce the agreed method rather than becoming the place where unresolved policy questions are hidden.
Implementation checklist#
- Define event, expected state and impact.
- Preserve evidence and build chronology.
- Identify immediate and contributing factors.
- Review control design, operation and dependencies.
- Apply proportionate RCA methods.
- Test alternative causes and causal link.
- Map actions to validated causes.
- Approve RCA and update risk records.
- Monitor recurrence and action effectiveness.
How Vilfora ERM can support the process#
Vilfora's Investigation and RCA workspace can record chronology, evidence, cause categories, analysis and approvals. RCA conclusions connect to incident actions, the issue register and linked risks and controls, enabling recurring cause analysis and tracking whether corrective actions prevent recurrence.
Suggested product screenshot: Vilfora Incident Investigation and RCA showing chronology, contributing factors, root cause, control failure and action links.
The screenshot should use anonymised demonstration data and should not expose personal information, credentials, confidential client information or internal environment details. Use a clear crop that shows the relevant workflow, status indicators and drill-down structure. Add a short caption explaining the management decision supported by the screen rather than merely naming the menu.
Frequently asked questions#
Is five whys enough for every incident?#
No. It is useful for relatively simple causal chains but may oversimplify complex events. Material incidents may require causal mapping, barrier analysis, fishbone analysis or specialist technical investigation.
What is the difference between a contributing factor and root cause?#
A contributing factor increased the likelihood or impact, while a root cause is a deeper condition whose removal would materially reduce recurrence. Several root causes can exist.
Who should approve RCA?#
The incident owner and relevant specialist functions should review it, with independent risk or compliance challenge for material events. Approval authority should reflect severity and regulatory significance.
Related reading#
- Risk Incident Management Process: From Event Capture to Closure and Learning
- Operational Loss and Near-Miss Management: Capture Better Data and Prevent Recurrence
- Incident Analytics: How to Identify Emerging Risk Patterns Before They Escalate
- Issue and Action Management: How to Close Findings Effectively and Prevent Repeat Issues
Final perspective#
Root cause analysis for risk incidents improves controls only when it moves beyond symptoms and blame. Evidence, chronology, multiple causal layers and challenge support a credible conclusion. Actions should directly address the causal pathway and their effectiveness should be monitored through recurrence, control testing and updated risk indicators.





