Unit 08 · Chapter 4 · 10 min read

Risk incidents, containment, and learning

Respond to active harm and preserve an accurate account of events.

The loss graph bends upward at 2 a.m. The first job is to stop the bleeding without breaking the rest of the payment system. The second is to understand what happened well enough that the next team does not repeat it.

Declare the incident from impact

An incident can involve active fraud, a broken sanctions feed, duplicate payouts, missing notices, data exposure, or a liquidity shortfall. Define severity from customer harm, financial exposure, legal duties, service impact, and uncertainty.

Use a clear declaration process with an incident lead and decision roles. Preserve the initial evidence and assumptions. Early estimates will change; version them rather than treating the first number as permanent truth. A small confirmed issue with large unknown exposure may need more attention than its current loss count suggests.

An incident can exist before the service is fully unavailable. If a release approves payments without a required check, every endpoint may respond quickly while the platform creates new exposure. Declare incidents from customer, financial, security, and control impact as well as infrastructure health. A shared severity definition helps teams recognize this class of failure early.

The initial scope will often be uncertain. Record the known start time, suspected affected population, systems involved, and the evidence supporting each statement. Update these facts as the investigation improves. A confident estimate without a reproducible population query can send recovery teams toward the wrong accounts and make later reconciliation harder.

Declare the incident from impact — the flow
Declare the incident from impact Declare the incident from impact — the flow Follow the sequence. Assign command and response roles. Detect Identify the abnormal outcome Assess Estimate harm exposure and uncertainty Declare Assign command and response roles
  1. DetectIdentify the abnormal outcome
  2. AssessEstimate harm exposure and uncertainty
  3. DeclareAssign command and response roles
Follow the sequence. Assign command and response roles. Chapter sources · Open image
Declare the incident from impact — the distinction
Declare the incident from impact Declare the incident from impact — the distinction These concepts answer different questions. Read each definition in the context of the section. Confirmed loss Evidence supports the amount Potential exposure Additional impact remains possible
Confirmed loss
  • Evidence supports the amount
Potential exposure
  • Additional impact remains possible
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Incident scope
Declare the incident from impact Incident scope Fictional teaching record. Not based only on four cases. Incident scope Illustrative data; not a real customer record or a prescribed policy. Confirmed duplicates 4 Known affected payments Unreviewed batch 2000 Potential broader scope Severity uncertainty included Not based only on four cases Known loss may be only part of the event
Fictional educational excerpt / Not for execution

Incident scope

Illustrative data; not a real customer record or a prescribed policy.

  1. Confirmed duplicates4

    Known affected payments

  2. Unreviewed batch2000

    Potential broader scope

  3. Severityuncertainty included

    Not based only on four cases

Known loss may be only part of the event

Fictional teaching record. Not based only on four cases. Chapter sources · Open image
Declare the incident from impact — control and failure modes
Declare the incident from impact Declare the incident from impact — control and failure modes Known loss may be only part of the event. The branches show why alternative designs fail. Control design Include uncertainty in severity and scope. Known loss may be only part of the event. Failure mode 1 Wait for perfect facts before organizing. Harm can continue. avoid Failure mode 2 Use the first estimate forever. Evidence develops. avoid Failure mode 3 Assign several competing incident leads. Decision authority becomes unclear. avoid
Control design

Include uncertainty in severity and scope. Known loss may be only part of the event.

Failure mode 1avoid
Wait for perfect facts before organizing. Harm can continue.
Failure mode 2avoid
Use the first estimate forever. Evidence develops.
Failure mode 3avoid
Assign several competing incident leads. Decision authority becomes unclear.
Known loss may be only part of the event. The branches show why alternative designs fail. Chapter sources · Open image

Contain the harmful capability

Containment should target the path causing harm: pause a payout route, restrict a compromised credential, stop a defective rule, or hold work under the approved contingency. Avoid a broad shutdown when a narrower action can safely stop the problem.

Record the authority, time, scope, and expected effect of each action. Confirm that it worked using direct evidence. A kill switch that changes a dashboard flag but leaves workers active is not containment. Consider customer consequences and legal constraints while preserving the records needed for investigation.

Contain the harmful capability — the flow
Contain the harmful capability Contain the harmful capability — the flow Follow the sequence. Observe that the harmful effect stopped. Target Identify the harmful path Act Apply an authorized bounded restriction Confirm Observe that the harmful effect stopped
  1. TargetIdentify the harmful path
  2. ActApply an authorized bounded restriction
  3. ConfirmObserve that the harmful effect stopped
Follow the sequence. Observe that the harmful effect stopped. Chapter sources · Open image
Contain the harmful capability — the distinction
Contain the harmful capability Contain the harmful capability — the distinction These concepts answer different questions. Read each definition in the context of the section. Command issued Containment was requested Containment effective Evidence shows the path is stopped
Command issued
  • Containment was requested
Containment effective
  • Evidence shows the path is stopped
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Payout containment
Contain the harmful capability Payout containment Fictional teaching record. Verify the downstream effect. Payout containment Illustrative data; not a real customer record or a prescribed policy. Switch disabled Control command Worker releases still occurring Action ineffective Response stop actual release path Verify the downstream effect A control flag may not stop running workers
Fictional educational excerpt / Not for execution

Payout containment

Illustrative data; not a real customer record or a prescribed policy.

  1. Switchdisabled

    Control command

  2. Worker releasesstill occurring

    Action ineffective

  3. Responsestop actual release path

    Verify the downstream effect

A control flag may not stop running workers

Fictional teaching record. Verify the downstream effect. Chapter sources · Open image
Contain the harmful capability — control and failure modes
Contain the harmful capability Contain the harmful capability — control and failure modes A control flag may not stop running workers. The branches show why alternative designs fail. Control design Verify containment at the effect boundary. A control flag may not stop running workers. Failure mode 1 Assume the command succeeded. The action needs evidence. avoid Failure mode 2 Delete records to stop processing. That destroys the investigation trail. avoid Failure mode 3 Ignore customer obligations during a pause. They still require handling. avoid
Control design

Verify containment at the effect boundary. A control flag may not stop running workers.

Failure mode 1avoid
Assume the command succeeded. The action needs evidence.
Failure mode 2avoid
Delete records to stop processing. That destroys the investigation trail.
Failure mode 3avoid
Ignore customer obligations during a pause. They still require handling.
A control flag may not stop running workers. The branches show why alternative designs fail. Chapter sources · Open image

Maintain a decision and evidence log

An incident log records observations, decisions, owners, timestamps, and evidence references. Keep facts separate from hypotheses. Use a common time basis and preserve source timestamps when they differ.

Record why a decision was reasonable with the information available then. This reduces hindsight distortion during review. Limit sensitive content to the appropriate access group and use references in broad coordination channels. The log should support handoff when responders change shifts and should make unresolved questions visible.

Maintain a decision and evidence log — the flow
Maintain a decision and evidence log Maintain a decision and evidence log — the flow Follow the sequence. Show unresolved work and next owners. Observe Record facts with sources Decide Capture rationale authority and time Handoff Show unresolved work and next owners
  1. ObserveRecord facts with sources
  2. DecideCapture rationale authority and time
  3. HandoffShow unresolved work and next owners
Follow the sequence. Show unresolved work and next owners. Chapter sources · Open image
Maintain a decision and evidence log — the distinction
Maintain a decision and evidence log Maintain a decision and evidence log — the distinction These concepts answer different questions. Read each definition in the context of the section. Hypothesis Possible explanation under investigation Established fact Supported observation or confirmed result
Hypothesis
  • Possible explanation under investigation
Established fact
  • Supported observation or confirmed result
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Incident log
Maintain a decision and evidence log Incident log Fictional teaching record. Action with evidence. Incident log Illustrative data; not a real customer record or a prescribed policy. 02:10 duplicate releases observed Fact 02:15 retry defect suspected Hypothesis 02:20 route paused and verified Action with evidence The record must show what was known when
Fictional educational excerpt / Not for execution

Incident log

Illustrative data; not a real customer record or a prescribed policy.

  1. 02:10duplicate releases observed

    Fact

  2. 02:15retry defect suspected

    Hypothesis

  3. 02:20route paused and verified

    Action with evidence

The record must show what was known when

Fictional teaching record. Action with evidence. Chapter sources · Open image
Maintain a decision and evidence log — control and failure modes
Maintain a decision and evidence log Maintain a decision and evidence log — control and failure modes The record must show what was known when. The branches show why alternative designs fail. Control design Version facts hypotheses and decisions. The record must show what was known when. Failure mode 1 Rewrite early notes as if the cause was obvious. That loses the real decision context. avoid Failure mode 2 Paste secrets into broad channels. Coordination does not remove access limits. avoid Failure mode 3 Handoff without open questions. The next team can miss critical work. avoid
Control design

Version facts hypotheses and decisions. The record must show what was known when.

Failure mode 1avoid
Rewrite early notes as if the cause was obvious. That loses the real decision context.
Failure mode 2avoid
Paste secrets into broad channels. Coordination does not remove access limits.
Failure mode 3avoid
Handoff without open questions. The next team can miss critical work.
The record must show what was known when. The branches show why alternative designs fail. Chapter sources · Open image

Recover with a reconciled population

Recovery should restore safe service and resolve the affected obligations. Identify all potentially affected records with a reproducible query, confirm impact, and assign the appropriate remediation. Keep uncertain records in scope until resolved.

Use idempotent correction jobs and independent reconciliation. A refund or reversal batch can itself create duplicate value if retried poorly. Resume service under defined guardrails and monitor the original failure mode. The incident is not complete merely because the graph returns to normal while historical customers remain affected.

Recovery is complete when the affected population has a defined and verified outcome. Some payments may need replay, some may already have succeeded, and some may require manual resolution or customer communication. Use stable identifiers to assign each item to a resolution state and verify totals against independent records. Do not replay an entire time range solely because the service was degraded during it. An outage window is a useful investigation boundary, but it does not prove that every action within the window failed.

Recover with a reconciled population — the flow
Recover with a reconciled population Recover with a reconciled population — the flow Follow the sequence. Verify safe service and remaining obligations. Population Identify and confirm affected records Remediate Execute controlled corrections Resume Verify safe service and remaining obligations
  1. PopulationIdentify and confirm affected records
  2. RemediateExecute controlled corrections
  3. ResumeVerify safe service and remaining obligations
Follow the sequence. Verify safe service and remaining obligations. Chapter sources · Open image
Recover with a reconciled population — the distinction
Recover with a reconciled population Recover with a reconciled population — the distinction These concepts answer different questions. Read each definition in the context of the section. Service recovery New activity works again Impact resolution Affected historical records have supported outcomes
Service recovery
  • New activity works again
Impact resolution
  • Affected historical records have supported outcomes
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Recovery population
Recover with a reconciled population Recovery population Fictional teaching record. Incident work remains. Recovery population Illustrative data; not a real customer record or a prescribed policy. Confirmed affected 180 Reviewed scope Corrected 176 Completed actions Unresolved 4 Incident work remains Restored service does not resolve past impact
Fictional educational excerpt / Not for execution

Recovery population

Illustrative data; not a real customer record or a prescribed policy.

  1. Confirmed affected180

    Reviewed scope

  2. Corrected176

    Completed actions

  3. Unresolved4

    Incident work remains

Restored service does not resolve past impact

Fictional teaching record. Incident work remains. Chapter sources · Open image
Recover with a reconciled population — control and failure modes
Recover with a reconciled population Recover with a reconciled population — control and failure modes Restored service does not resolve past impact. The branches show why alternative designs fail. Control design Reconcile remediation against the affected population. Restored service does not resolve past impact. Failure mode 1 Run unkeyed correction scripts repeatedly. They can create new duplicates. avoid Failure mode 2 Drop uncertain records from the count. The scope becomes falsely small. avoid Failure mode 3 Close at the first normal chart. Historical obligations may remain. avoid
Control design

Reconcile remediation against the affected population. Restored service does not resolve past impact.

Failure mode 1avoid
Run unkeyed correction scripts repeatedly. They can create new duplicates.
Failure mode 2avoid
Drop uncertain records from the count. The scope becomes falsely small.
Failure mode 3avoid
Close at the first normal chart. Historical obligations may remain.
Restored service does not resolve past impact. The branches show why alternative designs fail. Chapter sources · Open image

Learn from the system that allowed the event

A useful incident review explains the trigger, contributing conditions, detection, response, and customer impact. Avoid stopping at a person made a mistake. Examine why the system permitted the error and why existing controls did not catch it sooner.

Choose corrective actions that change the failure path and can be verified. Assign owners, dates, and acceptance evidence. Track completion and retest important controls. A long list of reminders can be less useful than one enforced invariant. The review should improve the system’s ability to prevent, detect, or contain the next similar event.

Learn from the system that allowed the event — the flow
Learn from the system that allowed the event Learn from the system that allowed the event — the flow Follow the sequence. Retest the relevant failure path. Explain Trace trigger and contributing conditions Repair Choose enforceable changes Verify Retest the relevant failure path
  1. ExplainTrace trigger and contributing conditions
  2. RepairChoose enforceable changes
  3. VerifyRetest the relevant failure path
Follow the sequence. Retest the relevant failure path. Chapter sources · Open image
Learn from the system that allowed the event — the distinction
Learn from the system that allowed the event Learn from the system that allowed the event — the distinction These concepts answer different questions. Read each definition in the context of the section. Individual action What a person did System condition Why that action could cause or escape harm
Individual action
  • What a person did
System condition
  • Why that action could cause or escape harm
These concepts answer different questions. Read each definition in the context of the section. Chapter sources · Open image
Post-incident action
Learn from the system that allowed the event Post-incident action Fictional teaching record. Verifiable acceptance. Post-incident action Illustrative data; not a real customer record or a prescribed policy. Cause duplicate posting allowed Missing invariant Fix unique operation constraint Enforced boundary Evidence retry fault test passes Verifiable acceptance Lessons should change behavior in the system
Fictional educational excerpt / Not for execution

Post-incident action

Illustrative data; not a real customer record or a prescribed policy.

  1. Causeduplicate posting allowed

    Missing invariant

  2. Fixunique operation constraint

    Enforced boundary

  3. Evidenceretry fault test passes

    Verifiable acceptance

Lessons should change behavior in the system

Fictional teaching record. Verifiable acceptance. Chapter sources · Open image
Learn from the system that allowed the event — control and failure modes
Learn from the system that allowed the event Learn from the system that allowed the event — control and failure modes Lessons should change behavior in the system. The branches show why alternative designs fail. Control design Choose repairs with observable acceptance evidence. Lessons should change behavior in the system. Failure mode 1 End at human error. Contributing design conditions remain. avoid Failure mode 2 Assign reminders without owners. The work may not happen. avoid Failure mode 3 Close after code review alone. The failure path still needs verification. avoid
Control design

Choose repairs with observable acceptance evidence. Lessons should change behavior in the system.

Failure mode 1avoid
End at human error. Contributing design conditions remain.
Failure mode 2avoid
Assign reminders without owners. The work may not happen.
Failure mode 3avoid
Close after code review alone. The failure path still needs verification.
Lessons should change behavior in the system. The branches show why alternative designs fail. Chapter sources · Open image

Chapter connections

This chapter builds on Operational resilience and failure design. Continue with A complete risk system: the Lantern case to follow the next part of the system. Use the glossary for terminology and risk mathematics for formulas and worked calculations.

Sources

Reviewed 2026-09-17
  1. NIST: Cybersecurity Framework
  2. Google SRE: handling overload
  3. Federal Reserve SR 23-4: third-party relationships