Work
The technical half
Case studies, the systems behind them in 3D, the incidents they produced, and analysis of other people's outages. All of it cross-linked — a case study links to its architecture and its incidents, and back again.
- Case studies10
10 pieces of work, written up
What the problem was, the order things had to happen in, the commands, and what the numbers actually moved.
- ·Cloud Cost Reduction Programme
- ·CI Quality Gate at Scale
- ·Self-Service Infrastructure Provisioning
- ·RAG Knowledge Assistant for Ops
OPEN ↗
- Architecture5
5 systems, rendered in 3D
Orbit, zoom and pan around the gate, the mesh, the self-healing loop and the platform.
- ·Production Security
- ·PR Quality Gate
- ·Zero-Trust Mesh
- ·Self-Healing Loop
- ·Self-Service Platform
OPEN ↗
- Field notes7
7 incidents and their fixes
What broke, why it broke, and what actually fixed it — each one rendered failing, then recovering.
- ·connection reset by peer, immediately after enabling STRICT mTLS
- ·APM init container stuck in ImagePullBackOff across several services
- ·Dashboards green, services listed as onboarded, no data arriving
- ·Trace ingestion stopped for every tenant at once
OPEN ↗
- Postmortems2
Public outages, read closely
Analysis of published incident reports from companies operating at a scale most of us only read about.
- ·Roblox — 73 hours
- ·Cloudflare — 27 minutes
OPEN ↗