At 9:47 AM on a Tuesday, every business application stops working simultaneously. Email authentication fails. The CRM returns error messages. The customer portal displays "unable to authenticate user." VPN connections drop. File shares become inaccessible. The help desk receives 200 calls in ten minutes, all reporting the same pattern: applications are running, but no one can log in. A certificate expired in the identity provider (IdP) at 9:45 AM, and every federated application lost the ability to verify user authentication. The business discovers it has an identity availability dependency only when the outage exposes it.
Identity outages expose a critical gap between business continuity planning and security operations. Business continuity teams plan application recovery with RTOs and RPOs, while security teams manage the IdP that every application depends on. The two disciplines rarely intersect — BCP teams treat authentication like network connectivity, assuming it will be available during recovery. When the IdP fails, application restoration procedures cannot restore application access because no user can authenticate to the recovered application. The consequence is runbooks that restore applications but not business function.
This is fundamentally an availability problem. Security frameworks organize around confidentiality, integrity, and availability — and identity infrastructure has historically been evaluated through the lens of confidentiality and integrity. Whether credentials are protected, whether authentication is strong, whether access is appropriately scoped. Availability receives less attention because identity systems are typically stable. That stability creates a false assumption: that availability is guaranteed. When the assumption breaks, it breaks across every application simultaneously.
What had changed
Two structural changes transformed identity from a security concern into a business continuity dependency. First, consolidation expanded the blast radius. Organizations that once maintained separate authentication for email, VPN, CRM, and file shares now route everything through a single IdP. This consolidation eliminated password management overhead and improved security policy enforcement, but created a hard dependency relationship that most organizations have not mapped explicitly.
Second, application design changed. Modern applications delegate authentication entirely to the IdP through SAML, OAuth, or OIDC federation. There is no local authentication fallback by design — applications verify tokens issued by the IdP rather than maintaining their own credential stores. When applications outsource authentication completely, they inherit the IdP's availability characteristics as an application dependency.
The combination creates a new failure mode: single points of authentication failure that affect multiple business functions simultaneously. An identity outage is not a security incident — it is an availability incident that affects every application simultaneously. The security team owns the system that fails, but the business continuity team owns the recovery objectives that cannot be met.
Why the current model fails
Three structural problems prevent effective identity availability planning. The planning gap appears first. The runbook for restoring the customer portal includes database recovery, load balancer configuration, and SSL certificate validation. It does not include "verify IdP availability" before declaring the application restored. BCP teams design application recovery assuming authentication works. When the assumption fails, every application recovery plan becomes invalid.
Dependency mapping does not exist. Most organizations cannot answer the question "what breaks if the IdP is unavailable for 30 minutes?" until something actually breaks it. Application owners understand their database dependencies and network requirements. They do not map authentication dependencies because authentication was traditionally local to each application. Discovery happens during outages, not during planning.
Break-glass access remains theoretical. Break-glass emergency access procedures exist on paper but lack operational muscle memory. The emergency accounts exist. The escalation procedures exist. The coordination between security and application teams during identity outages does not exist. Testing break-glass procedures requires deliberately breaking identity services — an exercise most organizations avoid.
Evidence and synthesis
Certificate expiration provides the canonical example of identity availability failure. IdP certificates expire on a scheduled date with 30-90 days' advance notice in monitoring systems. When the certificate expires, the IdP cannot sign new authentication tokens, and applications cannot verify existing tokens. The incident commander faces dozens of seemingly unrelated application outages that share the same invisible root cause.
Certificate expiration incidents follow a predictable pattern: applications fail authentication silently rather than displaying certificate error messages, making the root cause non-obvious during initial triage. Security teams receive alerts about authentication failures, while application teams receive alerts about user access problems, and the two sets of alerts are not correlated until someone maps the timing.
The failure mode reveals the dependency mapping gap. Applications that worked perfectly before 9:45 AM stop working after 9:45 AM, but application logs show normal operation. Application health checks may continue to pass because they test application functionality, not user authentication. Operations teams troubleshoot application issues while security teams troubleshoot identity issues, often without coordinating until the outage duration triggers escalation procedures.
Consequences
Identity outages expose operational assumptions that work until they fail catastrophically. Standard application restoration assumes authentication is available. Organizations measure restoration time from when the application responds to health checks, not from when users can actually work.
Break-glass procedures fail under real outage conditions because they require coordination across teams that do not work together regularly. Emergency access activation requires security team approval, but the security team may be focused on determining whether the identity failure represents a security incident. The procedural gap: emergency access requires cross-team coordination that does not happen during normal operations.
The business impact compounds during extended outages. Users cannot access email to receive outage communications. Support teams cannot access ticketing systems to track incident response. Management cannot access dashboards to assess business impact. The identity outage becomes self-reinforcing because the tools needed to manage the outage require authentication.
What happens next
Two converging forces are making identity availability a regulatory and architectural imperative. Financial services organizations subject to DORA requirements must demonstrate operational resilience for critical business functions, including the identity systems that enable access to those functions. Healthcare organizations under HIPAA availability requirements cannot treat identity failures as out-of-scope when those failures prevent access to patient care systems. Compliance frameworks are beginning to evaluate identity infrastructure availability explicitly, not just security.
The architectural direction increases the stakes. Zero-trust architectures position identity verification as the security boundary for every access decision. Organizations building zero-trust implementations focus on the security benefits without mapping the availability dependencies they are creating. Identity becomes the control plane for security policy enforcement, which means identity availability becomes the availability of the security perimeter itself.
The structural question organizations must answer: if identity is the control plane, what is its recovery time objective? The business continuity planning gap that created certificate expiration outages will create zero-trust availability failures unless identity moves from security infrastructure to business continuity infrastructure. The question is not whether identity will be included in business continuity planning, but how long organizations will wait before the inclusion becomes mandatory.
Sources:
- CyberArk: Machine Identity Sprawl and Expired Certificates
- HHS.gov: Summary of the HIPAA Security Rule
- IBM: Digital Operational Resilience Act
- IBM: Zero Trust Implementation
- Microsoft Learn: Azure Reliability Incident Response
- Microsoft Learn: Azure Foundry Disaster Recovery
- Microsoft Learn: Emergency Access in Entra ID
- Microsoft Learn: Federated Single Sign-On Certificate Management
- Microsoft Learn: Basic Authentication Deprecation
- Microsoft Learn: MFA Lockout Resolution
- Microsoft Learn: Zero Trust Overview
- MITRE ATT&CK: T0815 - Spoof Reporting Message
- MITRE ATT&CK: T0826 - Loss of Availability
- MITRE ATT&CK: T1556 - Modify Authentication Process
- MITRE ATT&CK: T1562 - Impair Defenses
- MITRE ATT&CK: Enterprise Techniques
- NIST SP 800-63-3: Digital Identity Guidelines
- NIST SP 800-63C: Digital Identity Guidelines
- NIST SP 1800-16: Securing Web Transactions
- attack.mitre.org
- www.cyberark.com
