Field Notes

Where systems fail
and what it costs.

Short observations for practitioners and deep-dive analyses of the failures that define production AI. Every post is anchored to a real incident, a verifiable finding, and one thing you can act on before the next sprint planning.

Architecture
The CAB approved every change. The system still went down 16 times.
June 2026
Change approval is not the same as change safety. DORA's four-year research programme established exactly what the difference is — and why the organisations with the most rigorous CAB processes often have the highest change failure rates.
4 min read
Architecture
Your deployment pipeline has five stages. Your AI model has one.
June 2026
The same organisation that would never ship an application without automated quality gates is deploying models with a single stage: approved-by-data-scientist to production. A Canadian tribunal just ruled on what that asymmetry costs.
4 min read
Architecture
Your architecture diagram is correct. Your architecture is not.
June 2026
87% of AI projects never reach production. The failure is not in the code. It is in what the design stage chose not to represent — and the EU AI Act has made that omission a structural compliance obligation.
4 min read
Architecture
The Equifax breach. The patch was in the backlog.
June 2026
The Apache Struts vulnerability was publicly known and patchable for 78 days before the breach began. A known vulnerability you have not acted on is not a risk. It is a decision. The FTC's consent order names what that decision costs.
4 min read
Architecture
Your monitoring dashboard monitors the wrong layer.
June 2026
In January 2024, DPD's AI chatbot began swearing at customers and writing poetry criticising the company. Infrastructure monitoring was green throughout. The output layer had no monitoring at all.
4 min read
Architecture
Conway's Law doesn't care about your AI strategy.
June 2026
Your AI strategy describes a unified platform. Your org chart will build six siloed pipelines. The Target breach of 2013 is the most precise documentation of what happens when architecture and communication structure diverge.
4 min read
ML Engineering
Explainable to you is not explainable to the decision-maker.
June 2026
26,000 Dutch families were wrongly flagged as fraudulent by an algorithm. The model was documented. It was not explainable in the only sense that matters: the person acting on its output could not interpret it.
4 min read
ML Engineering
You solved the right problem. You measured the wrong thing.
June 2026
The model was performing exactly as trained. The proxy metric had diverged from the business outcome it was meant to represent. Standard drift detection confirmed nothing was wrong. That is the failure mode nobody instruments for.
4 min read
ML Engineering
Too complex to hand off. So it never left the team that built it.
June 2026
Google Flu Trends was celebrated as outperforming the CDC. Within five years it was overestimating flu prevalence by a factor of two and was quietly retired. The post-mortem is a description of what happens when a model cannot survive the team that built it.
4 min read
ML Engineering
Your model passed every evaluation. The team built the eval set.
June 2026
When the team that trains a model also builds the evaluation set, the evaluation tests what the team thought of — not what production will encounter. The MHRA's 2022 guidance establishes why independent evaluation is a structural requirement, not a process preference.
4 min read
ML Engineering
Your feature pipeline has never been tested in production conditions.
June 2026
A LinkedIn algorithm update sent unexpected recommendation emails to millions of users — not because the algorithm was wrong, but because the feature pipeline had never been tested under full social graph density at production scale.
4 min read
ML Engineering
You version your model. You don't version the decision.
June 2026
Model versioning is standard practice. Decision versioning is not. The Dutch Belastingdienst case established what happens when individual decisions cannot be reconstructed after the fact — and Article 12 of the EU AI Act names the structural response.
4 min read
Practice Leadership
"We'll know it when we see it." That project is in the 87%.
June 2026
The UK Home Office's Visa Streaming Tool was suspended after an independent review found its success criteria had never been formally defined. The team knew what they wanted. Nobody had written it down. That distinction is everything.
4 min read
Practice Leadership
Your AI project has a sponsor. It doesn't have an owner.
June 2026
The Post Office Horizon scandal persisted for over two decades in part because nobody with decision-making authority was accountable for the system's correctness in operation. Sponsorship and ownership are not the same role.
4 min read
Practice Leadership
The team delivered exactly what was asked. The business couldn't use it.
June 2026
IBM Watson for Oncology was trained on hypothetical cases from a single US cancer centre. When deployed internationally, its recommendations conflicted with local guidelines. The technical delivery was complete. The deployment context had never been specified.
4 min read
Practice Leadership
100% utilisation. That's why nothing is getting delivered.
June 2026
Healthcare.gov launched in 2013 with every team fully committed and no capacity reserved for integration testing. The GAO's post-mortem and Accelerate's research quantify exactly what full utilisation does to delivery throughput.
4 min read
Practice Leadership
Your AI vendor gave you an SLA. It doesn't cover what will actually fail.
June 2026
Vendor SLAs cover uptime and latency. They do not cover the output — what the model says to your customers, whether it is correct, or who is liable when it is not. The Air Canada tribunal ruling and Article 25 make the same point from different directions.
4 min read
Practice Leadership
Delivered on time. Deadline set before anyone looked at the data.
June 2026
The Australian National Audit Office's 2022 review of the ATO's data analytics programme found timelines committed before data readiness was assessed — consistently. Delivery on time. Not what the programme needed.
4 min read

No notes in this category yet. Check back soon, or read everything under All.

Stay informed

New Field Notes
by email.

These notes are published when there is something worth saying — not to a schedule, not to fill a content calendar. If the observation is precise and the data supports it, it gets written. If it does not, it does not. To receive new Field Notes directly, send an email to the address below.

Subject line: Field Notes