Insights
Hard truths about
platforms, cloud, and AI
Short, direct takes on the problems that cause enterprise systems to fail. No fluff. No theory. Just signal.
Your agent is not broken. Your platform was never built to run non-deterministic work.
Gartner projects 40% of enterprise agentic AI initiatives will be cancelled by 2027. Industry reporting puts the 'never reaches production' rate at 88%. The model is not the problem. The platform underneath is.
Read the insightI joined this week. Only one person could deploy to France. That is not a deployment process.
I walked into a platform this week where production deployment to one region depended on a single engineer being available. Ansible playbooks. No CI gates. No automated promotion. Here is the failure mode and how I shipped a fix in a week.
Read moreI logged in this week. Datadog was down for two weeks. Nobody knew.
I joined a client this week and went to look at the application logs. There were none. The Datadog agent had been broken for weeks. Nobody had noticed. This is what monitoring theatre looks like in 2026.
Read moreMulti-cloud is a hedge against your own bad decisions. Most companies should not.
Multi-cloud is sold as resilience. It is usually overhead. Here is when it actually pays off, when it is just expensive, and the question I ask before I let a client commit.
Read moreOn-call rotation is your real architecture document. Read it.
Architecture diagrams lie. The on-call rotation does not. The most-paged service is your most-brittle dependency, and it is rarely the one in the strategy deck.
Read moreSLOs are being gamed. Here is what your team is hiding.
Your team is reporting 99.95% SLO. Customers are complaining. Both can be true. Five ways SLOs get quietly gamed, and how to ground them in customer reality.
Read moreYour CISO is right about AI. The fix is platform engineering, not policy.
Every AI rollout I have worked on has a moment where the CISO becomes difficult. The CISO is right. The policy approach they have to use is wrong. Here is what fixes it.
Read moreThe migration nobody finishes. Four I have rescued.
I have walked into four enterprise migrations stuck at 60% completion for years. Same pattern every time. Here is what the last 40% actually looks like.
Read moreYour observability stack is theatre. Here are the three signals you are missing.
Most observability stacks measure what is easy, not what is real. Three signals matter. Most teams capture none of them.
Read moreI refuse AI projects without 30 days of platform discovery. Three CTOs taught me why.
Every AI project I have rescued had the same root cause. The platform was not ready. The team did not know. The vendor did not care. I will not take that work without 30 days of discovery first.
Read moreStop hiring AI Engineers. Hire platform engineers who can read papers.
The AI Engineer job description is wrong. It is asking for the wrong skills, attracting the wrong candidates, and producing teams that cannot ship to production. Here is the role you actually need.
Read moreEvery enterprise has an AI strategy. Almost none have an AI operations plan.
The board approved your AI strategy. But nobody planned how to run AI systems in production at 2am when the model starts returning garbage. That gap is where the next outage is hiding.
Read moreYour AI compliance audit will fail. Here is why.
Most organisations cannot prove they are doing anything right with AI. They lack the infrastructure for data lineage, decision logging, and explainability.
Read moreWhy every AI incident becomes a cross-team incident
Three teams own different parts of the AI pipeline. Each team's monitoring shows green. The failure exists in the interaction between components.
Read moreThe real cost of AI is not the model. It is the data pipeline.
Every AI business case focuses on model costs. They are also the minority of the total cost. The data pipeline is typically 60-70%.
Read moreAI rollbacks are harder than you think
Rolling back a model is not like rolling back code. The output distribution changes, and dependent state becomes inconsistent.
Read moreMost AI monitoring is just uptime monitoring with a new label
Your AI monitoring checks that the service is responding. It does not check that the service is correct. That gap is where incidents hide for weeks.
Read moreYour Kubernetes cluster was not designed for GPU workloads
Standard K8s clusters are optimised for stateless, CPU-bound workloads. AI inference breaks all of those assumptions.
Read morePrompt injection is an infrastructure problem, not an AI problem
If your defense is a regex in your application code, you are playing a game you cannot win. AI security needs to live at the platform layer.
Read morePlatform engineers will own AI governance by 2027
The governance problems that cause production incidents are infrastructure problems. Platform teams already know how to solve them.
Read moreWhy your AI fallback is more dangerous than the failure
When AI fails hard, someone gets paged. When it fails softly with a fallback, the system keeps serving bad results and nobody notices.
Read moreAI workloads are hiding in your cloud bill
Nobody knows what AI actually costs because inference runs on shared compute with no attribution. That is a platform architecture problem.
Read moreYour AI pipeline has no owner. That is the real risk.
The ML team built it. The data team feeds it. The platform team hosts it. Nobody owns the whole thing. That gap is where incidents hide.
Read moreThe 3 infrastructure failures every RAG system hits
Vector DB scaling, embedding pipeline throughput, and retrieval quality degradation. The model is usually fine. The platform underneath is not.
Read moreMost AI guardrails only protect the demo
The guardrails most teams put around AI systems are tested against friendly inputs. Production is none of those things.
Read moreField notes, in your inbox
One short, opinionated piece per fortnight on platform engineering, cloud, and making AI work in production. No spam. Unsubscribe anytime.