Incident reports

When something breaks, it gets written down.

Every report here is a real incident on Walter Infrastructure's own systems, documented exactly as a client incident would be, then sanitized for publication. Client incidents are never published.

Published reports

Real failures, documented end to end.

2026-07-31 Storage Array Multi-Device Drop and Recovery Severity 1 Restored

Four of twelve devices in a RAID 6 array stopped responding at the same instant, exceeding the array's two-device tolerance and taking a 26.6 TB volume offline. The failure mechanism was traced to a hang in the SATA fan-out stage rather than to device failure. Service restored in 32 minutes, every dataset present and readable.

Ref: WI-2026-0731-01 Time to restore: 32 min Data loss: None found 19 pages
What these are

A report is not a status update.

Infrastructure work should not disappear into one person's memory, an email thread, or a terminal history. When something significant happens on a system I am responsible for, it leaves a durable record: what happened, what the evidence showed, why each decision was made, what was ruled out, what is still uncertain, and what the next engineer should do about it.

A status update expires when service comes back. The record does not. It keeps working after the incident closes: it stops a future investigation from re-running tests that already produced an answer, it writes down the platform quirks and traps that otherwise live in one engineer's head, and it gives whoever responds to a recurrence a tested starting point instead of a blank screen.

Who they serve

One document, three readers.

The person who runs the business gets the first page: what stopped, who felt it, how long restoration took, what the evidence says about the data, and what is still open. Restoration and resolution are kept separate on purpose. A system can be back online while risk is still standing, and the report says so plainly.

The technical reader gets everything behind it: the architecture, the timeline, the measurements, and the reasoning, including the working theories that testing killed. An analysis that lists only the correct answers gives a reader no way to judge it, so the dead ends stay in.

Whoever touches the system next gets the record itself. That may be me, the client's own staff, another consultant, or someone responding to a recurrence years later. Your infrastructure knowledge should belong to your organization, not sit in a consultant's head, so the report is written to be useful without its author in the room.

Why sanitized

Published because it is mine.

A client's report is not sanitized. Clients receive the full confidential record, unredacted, and it is never published. This one is different for a single reason: the incident happened on Walter Infrastructure's own systems, so publishing it is my call to make, and I made it so you can see exactly what you get as a client. The sanitization is the price of publishing: host names, network addresses, serial numbers, and the nature of the stored data are generalized, and the technical content is otherwise left intact.