Business Continuity & Risk

Set Action Thresholds for an External Service Outage

Use business impact, workaround capacity, and evidence checkpoints to decide when an external outage requires action instead of repeatedly waiting for recovery.

FIELD GUIDEPractical guide

Built for practical decisions, implementation, and review.

Overview

Set an outage action threshold at the point where waiting would leave too little time to protect an important business outcome. Use the affected function, the rate at which harm grows, and the time needed to activate an alternative. A provider's estimated recovery time is useful evidence, but it should not be the only trigger.

Teams often lose options by treating every update as a reason to wait another hour. The provider may be making a reasonable estimate while the customer's own deadlines become steadily harder to meet. A threshold gives the business a decision rule it can apply even when external information remains uncertain.

Define the dependency at the business level

Name the external service and the work that cannot proceed without it. “The booking provider is unavailable” is a technical description. “Customers cannot reserve tomorrow's appointments, and staff cannot see recent changes” identifies the operating consequences.

Separate affected functions. The same outage might stop new transactions while existing records remain readable. A workaround suitable for taking inquiries may be unsuitable for confirming bookings or accepting payment.

CISA's infrastructure dependency primer describes how dependencies can produce cascading effects. In a business plan, trace those effects one step further: which customer promise, staff activity, or supplier action depends on the unavailable function?

Record the dependency and consequence in the risk register, then keep the operational thresholds in a place staff can use during an incident.

Identify when impact changes, rather than assuming it is linear

Some outages cause a steady backlog. Others become much more damaging at a specific cutoff. A missed dispatch window, payroll submission deadline, or customer event can make the fifth hour very different from the first four.

List the relevant clocks. These may include the next required output, the point at which records become stale, the maximum manageable queue, or the expiration of a temporary arrangement.

Do not equate inconvenience with equal priority. A reporting dashboard may be unavailable for several hours without affecting today's service, while a short interruption in an approval route could block a time-sensitive delivery.

The CISA service continuity guide uses business impact analysis to establish continuity needs and recovery objectives. Your threshold should translate that analysis into a decision that an incident owner can make.

Work backward from the outcome

Suppose an illustrative business needs accepted orders ready for a carrier collection at 4 p.m. Its alternative process takes two hours to activate and prepare the priority orders. Waiting until 3 p.m. to decide cannot protect that collection even if the alternative works perfectly.

The latest decision time would therefore need to be earlier than the required outcome by at least the activation and execution time, with a suitable allowance for uncertainty. The exact margin should reflect tested performance rather than a universal rule.

Distinguish three moments: when the problem is recognized, when a decision must be made, and when the consequence occurs. Collapsing them into one “four-hour outage trigger” can hide the preparation time.

If the alternative has never been tested, treat its estimated activation time cautiously. A threshold built on an optimistic workaround estimate may be too late even when followed precisely.

Use separate thresholds for separate actions

Not every threshold should trigger a full switch. Early actions might include assigning an incident owner, pausing new promises, checking the queue, or preparing an alternative. Later actions might activate the workaround or suspend affected transactions.

A practical sequence can be described as awareness, preparation, commitment, and reassessment. The labels are optional; the distinction between reversible preparation and consequential switching is useful.

For example, staff can prepare a customer notice before deciding to publish it. They can confirm that an alternative provider is available before placing an order. Early preparation preserves options without forcing an unnecessary change.

Assign authority for each action. A front-line worker may be able to pause a promise but not authorize spending. If the threshold requires an unavailable approver, it is not operationally complete.

Include capacity and data quality

An outage workaround may process less work or use older information. Its usefulness depends on how long it can operate within those limits.

Estimate incoming work and processing capacity. If requests arrive faster than the workaround can handle them, define which requests receive attention and when the backlog requires another decision. Use the manual workaround throughput guide to ground this in observed performance.

Set a data-quality boundary. Staff may be able to record a request safely while being unable to confirm stock, availability, or account status. The threshold may therefore call for a narrower service rather than an attempt to recreate the unavailable system.

Do not let an alternative introduce uncontrolled duplicate transactions. Identify what can be committed, what remains provisional, and how it will later be reconciled. A workaround that creates uncertain obligations can extend the incident after the provider recovers.

Treat provider updates as evidence with a timestamp

Capture the provider's current status, affected components, stated estimate, and time of update. Distinguish confirmed restoration from an expected restoration window.

A provider announcement may cover only part of the service. Test the business function through the authorized route before assuming the incident is over. A status page marked healthy does not prove that your own queued transactions have completed.

When an estimate moves, compare it with the business threshold rather than simply resetting the waiting period. If the latest credible recovery window is beyond the time needed to activate the alternative, escalation should occur.

Also allow for uncertainty in the other direction. The provider may recover sooner than expected. The plan should explain how to stop preparation or safely reverse a partial switch without leaving conflicting records.

Make the decision record short enough to use

Record the current impact, evidence, threshold reached, action chosen, owner, and next review point. This can be a compact incident entry rather than a long report.

If the owner decides to wait beyond a planned threshold, record the reason and the accepted consequence. Perhaps new evidence makes recovery credible, or the alternative has become unavailable. An explicit exception is more useful than silently ignoring the rule.

Set the next review around the next meaningful change. Repeated meetings every few minutes can consume the people needed to respond; a long gap can allow the decision window to pass. Choose a cadence that matches the operational clock.

Make the time zone explicit when teams or providers work in different regions. A cutoff recorded as 2 p.m. without a location can produce incompatible decisions, especially when an external update uses a different clock.

Keep customer-facing messages consistent with the decision. If new commitments are paused, all relevant channels should reflect that state.

Rehearse the awkward scenario

In a tabletop exercise, make the provider's estimate slip more than once. Then add a constraint: the backup coordinator is unavailable, the queue is larger than expected, or one part of the service recovers while another remains down.

Ask the team which threshold now applies and who can act. The exercise should reveal whether the rule survives imperfect information.

Use tabletop action tracking to close gaps in authority, timing, or evidence. Update thresholds when actual operating capacity or business deadlines change.

A good threshold does not predict the outage duration. It protects the time and authority needed to respond while there are still useful choices available.

References and examples

Primary sources and product examples used to ground this guide. Product links are editorial references, not endorsements.

Written and reviewed by

Smarter Business Results Editorial Team

We turn source research and operational questions into independent, practical frameworks. We do not invent product capabilities, credentials, or results.

Source review .

Search the library

What decision are you working through?

Try “automation,” “electronic signatures,” “modular home,” or “product feedback.”