Nicholas TongCrowdStrike Falcon
50 / 56
the night the security agent was the outage
I read a 48-digit number out loud four times that night, once per machine, off a printout that had been living in a filing cabinet for over a year. Someone at the MSP had laughed at me for making it. A spreadsheet of BitLocker recovery keys, on paper, printed single-sided because the only printer in the office with a working duplex unit had been broken since 2022. Paper. In a job about computers.
It was 11:09pm in Houston on Thursday, July 18, 2024, which was 04:09 UTC the next morning, and the on-call phone lit up. Then it lit up again. Windows machines across a chunk of our managed fleet were boot-looping, crash screen, reboot, crash screen, and the machines had one new thing in common that afternoon: a content update from CrowdStrike Falcon. The security agent had become the outage. By midnight we were driving to a clinic in Katy with a keyboard and a printed support article.
the part everyone gets wrong
The lazy lesson everyone walked away with is "CrowdStrike bad", and I don't buy it. I'll disclose the obvious first: I ran Falcon under an employer license at that job, client fleets included, so this is not a product I bought and it's not a purchase recommendation. What I can say from two years of living in its console is that its detections were among the best I've worked with. The night I'm describing happened to be its fault. Both things are true, and holding them together is the whole post.
Here's the actual lesson, and it isn't about CrowdStrike. Any kernel-level agent that auto-receives content is a change-management system you don't run. Falcon updates itself in two ways. There are software updates, which behave like software updates. And there is content: Rapid Response Content, the behavior definitions the sensor uses to recognize badness, delivered straight from the vendor's cloud without asking you, because attackers don't respect your patch window. That speed is the product. It's why a detection for a new technique could exist on your fleet hours after someone saw it in the wild. The same property, on the worst night of the year, put a bad file on millions of machines with no ring, no canary, and no pause for a human.
the mechanics of those seventy-eight minutes
The timeline from CrowdStrike's own technical write-up is short. The Channel File 291 update went out at 04:09 UTC on July 19 and was reverted at 05:27 UTC. Seventy-eight minutes, start to revert. Microsoft later put the affected count at about 8.5 million Windows devices, under one percent of Windows machines, which is the least comforting one percent in history, because those machines happened to sit in airlines, hospitals, broadcasters, and the back offices of everything else.
The root cause analysis, published August 6, is worth reading in the full version, because the executive summary phrases the bug backwards. In the body of the RCA it's clear: the IPC template type defined 21 input fields, the integration code supplied 20, and the Content Interpreter read out of bounds. One missing field in one validation check, in content, not code, and the sensor died at boot and took the operating system with it. A config file with an off-by-one. That's the part I want everyone to sit with: the thing that failed was data, validated by code that didn't check it.
The recovery was worse than the diagnosis. The fix was reboot into safe mode, delete one file, reboot, and safe mode on a BitLocker-protected machine asks for its recovery key, all 48 digits of it. Keys normally live in Active Directory or a cloud console. On clients where the domain controller was one of the boot-looping machines, the key was stored on a computer that was down, waiting to be read by a computer that couldn't boot. The only copy that helped that night was the one nobody was supposed to need: mine, on paper, in a cabinet.
what it cost, in hours
Per machine, the fix was five to thirty minutes of console time, and the machines were not remote-accessible, because a machine that won't boot has no agent and no agent means no remote access. So the hours multiplied by driving. My night ran to about seven hours and four recovery keys read aloud, twice each, because bored and stressed people mishear digits. One client's clinic opened two hours late. Somewhere in the middle of that, around 2am, I admitted to myself that I had never once walked the safe-mode-and-delete procedure on a test machine. I printed the spreadsheet after a ransomware job in 2021 and then filed it, and filing it felt like finishing. A recovery key you've never tested is a theory. Backups are not real until you have restored from them, and I had never restored from that spreadsheet until it was the only thing between a clinic and a closed sign.
The industry-scale bill came later. Delta's claims, gross negligence and computer trespass, were allowed to proceed in May 2025, with no trial date as of early 2026, and the shareholder suit was dismissed in January 2026 with leave to amend. I'm not going to characterize fault past what the filings say. What the filings say is that the outage was expensive enough to still be in court two years later.
the question to ask a vendor instead
CrowdStrike, to its credit in the RCA, committed to staged deployment rings and customer control over Rapid Response Content. Those are the correct words. So the question I now ask any endpoint vendor, before I look at a single detection benchmark, is this: how does your content ship, and who can stop it? Rings, canary populations, a hold switch, a rollback path, and whether a customer can pin a version. Detection rates are marketing. Rings are architecture.
Second: keep recovery keys somewhere that survives the domain dying. Paper counts. Cloud consoles count, if they're independent of the fleet. Test one key per year by walking the recovery screen, because untested recovery is a hope with numbers in it.
Third: assume the update pipeline is part of your attack surface. In September 2025, npm packages published by CrowdStrike were swept up by the Shai-Hulud worm, and CrowdStrike says those packages weren't used in the Falcon sensor. I believe them, and it doesn't matter, because the pattern is the point: the channel that updates your security tool is trusted by every machine it touches, and that channel should be on your risk register the way your domain admins are.
Nobody at the MSP has laughed at the printout since.
What I keep from that night is smaller and stranger: Falcon worked. In the months after, it kept catching the things it caught before, on the same consoles, at the same speeds. The agent that melted a Friday was still the best detector on the fleet. That's not a contradiction, and any security engineer who tells you the tool with the worst outage is the worst tool hasn't carried the pager for the tool with the best one.
Go print your recovery keys. Then test one.