Image Credits: Harun Ozalp / Anadolu / Getty Images & Crowdstrike logo

On July 19, 2024, businesses and services worldwide faced significant disruptions due to a major IT outage linked to CrowdStrike, a leading cybersecurity firm. This incident affected numerous industries, including airlines, banking, media, healthcare, and local government, leading to travel chaos and operational delays.

The Incident

The outage was caused by a faulty update from CrowdStrike that impacted systems running on Windows. The update led to widespread issues, including the infamous “blue screen of death” errors on numerous computers. CrowdStrike quickly acknowledged the problem, emphasizing that it was not due to a cyberattack but a software issue. They have since deployed a fix and are working with affected customers to restore normal operations.

Impact on Businesses and Services

The outage’s repercussions were extensive:

  • Airlines: Major hubs like Newark International and Berlin’s BER airport experienced delays, with passengers stranded and flights grounded. United Airlines, along with other top carriers like Delta and American Airlines, were significantly impacted.

“Our operations were heavily impacted by the CrowdStrike outage, causing delays and cancellations. We are working to restore normalcy as quickly as possible.” ~United Airlines

  • Banking and Financial Services: Large financial institutions such as JPMorgan Chase and Citibank reported service interruptions, impacting transactions and customer service.

“The IT outage disrupted our services, affecting transaction processing and customer service. We are collaborating with CrowdStrike to ensure full recovery.” ~JP Morgan Chase

  • Media and Retail: Companies like Target and media conglomerates experienced operational disruptions, highlighting the critical dependency on digital infrastructure.
  • Healthcare Services: Hospitals and clinics, including major networks like Mayo Clinic, faced interruptions in their IT systems, affecting patient care and administrative functions.
  • Local Government and Emergency Services: Vital services, including 9–1–1 dispatching, experienced outages, compromising emergency response capabilities in cities like New York and Los Angeles.

CrowdStrike’s Role and Response

CrowdStrike, founded in 2011 and headquartered in Austin, Texas, is renowned for its Falcon platform, which identifies and mitigates security threats. The company’s services are utilized by high-profile clients like Google, Amazon, and the U.S. government. Despite the outage, CrowdStrike’s proactive response and swift deployment of a fix demonstrate their commitment to resolving the issue and maintaining customer trust. However, the recovery process requires manually updating and restarting each affected computer with the patched update before systems can come back online. This complex and time-consuming task is expected to take several days to fully restore normal operations across all affected systems.

Stock photo of a United Airlines plane. Getty

Market Reactions

In the immediate aftermath, CrowdStrike’s stock plummeted nearly 12% in premarket trading. However, the company’s strong market presence and reputation for robust security solutions are likely to help it recover in the long run.

Lessons Learned

1. Significant Impact of Relatively Unknown Vendors: The CrowdStrike outage highlights how even relatively lesser-known vendors like SolarWinds and CDK Global can have significant, global impacts. These incidents stress the interconnected nature of modern digital infrastructures and the potential risks posed by any single point of failure. Additionally, the unknown and critical dependencies across ecosystems and supply chains can exacerbate these impacts.

2. Need for Robust Validation Processes: The incident underscores the inadequacy of current patch/update validation processes. The lack of a resilient testing mechanism before pushing updates to production can have massive, industry-wide impacts. This necessitates the development of new IEC/ISO-level procedures and guidance for updates. For this particular failed update, best practices were not followed, including pre-deployment testing and phased rollouts.

“The recent CrowdStrike update failure serves as a critical reminder of the importance of comprehensive testing and phased deployment strategies. The failure to adhere to these practices can lead to significant operational disruptions.” ~Paul Proctor, Distinguished VP Analyst, Gartner

3. Hesitancy in OT Asset Updates: Operational Technology (OT) asset owners often hesitate to update systems due to concerns about operational resilience and uptime. In their risk analysis, they can’t afford the possible downtime which would impact production and operations. This reluctance can leave systems vulnerable to security threats, emphasizing the need for a balance between security and operational stability.

Moving Forward

The CrowdStrike outage serves as a critical reminder of the complexities and risks inherent in modern IT ecosystems. Moving forward, the IT industry must prioritize the development and adherence to rigorous testing protocols and robust update procedures. This incident emphasizes the importance of comprehensive pre-deployment testing, phased rollouts, and clear rollback plans to prevent widespread disruptions. Additionally, fostering greater transparency and communication within supply chains and ecosystems can help identify and mitigate potential risks.

To support critical infrastructures, the industry needs to adopt and enforce standardized guidelines, such as those provided by IEC and ISO. These standards will help ensure that updates are thoroughly tested and validated, minimizing the risk of future outages. By focusing on these areas, the IT industry can better support the resilience and security of critical infrastructures worldwide.

For more detailed information, you can refer to the original articles on TechCrunch and Reuters.