Systems In The Wild

Architecture observations on complex distributed systems.
An adult periodical cicada with red eyes and orange-veined wings, resting on tree bark in a summer forest

Every Step Reported Success

Ten articles in this series produced nine named measurements. Not one of them measures a component. They measure distances. Between a policy that exists and a policy that arrives. Between a recovery time that was written and one that was observed. Between an action taking effect and the judgment about it arriving. That is not a stylistic habit. It is the finding, and it went unnoticed while it was being assembled, because each article had exactly one instance in front of it. ...

August 26, 2026 · 18 min · 3809 words · Andre Rocha
A leopard gecko on rocky arid ground at dusk, its banded tail intact and extended behind it, scrubland and low hills beyond

The Last Human Step

In this industry the word has already changed hands. An operator used to be a person. Today an operator is more likely to be a controller: software that watches a declared state, notices drift, and corrects it without being asked. The migration was so complete that the original meaning now requires a qualifier. The vocabulary vanished before the role did. ...

July 28, 2026 · 25 min · 5321 words · Andre Rocha
A cluster of acorn barnacles encrusting a coastal rock, with surf breaking behind

Data Sovereignty as an Architectural Constraint

Data sovereignty has long been treated as a compliance topic. The treatment is accurate but incomplete. Compliance frameworks document the commitment. Architecture determines whether the commitment can be kept. The distinction is structural. Compliance achievements can be documented, certified, and renewed annually. Architectural commitments are designed in and become harder to change with every service, dataset, and integration that depends on them. Organizations that maintain rigorous compliance and weak architectural sovereignty share a common position: the documentation passes audit, and the architecture cannot be moved. ...

July 7, 2026 · 13 min · 2574 words · Andre Rocha
Peacock mantis shrimp on a coral reef, full body with raptorial appendages, antennae, and stalked eyes visible

The Dashboard Illusion

Observability is described as understanding the system. It is detection. The distinction is not academic. It is the difference between knowing that a signal exists and knowing what the signal means about the platform that produced it. Detection has been industrialized over the past decade. Understanding has not. Most of the friction during incidents lives in the gap between them. This article is not an argument against observability. The detection capability the industry has built is real and valuable. The argument is that detection has been mistaken for comprehension, and that the conflation has a measurable cost. ...

June 16, 2026 · 10 min · 2089 words · Andre Rocha
Macro photograph of a tardigrade on wet moss

The DR Number Almost No One Records

Disaster recovery has three numbers. Almost no organization records all three. The first is the number written into the plan. The second is the number measured during exercises, if exercises happen. The third is the number observed during real incidents. The distance between them is the only metric that matters. It is also the metric that almost no one calculates. The Three States of D.R. Capability Disaster recovery capability exists in three forms simultaneously, and the three forms produce three different numbers. ...

May 28, 2026 · 10 min · 1944 words · Andre Rocha
Coral colony of many small polyps on one shared skeleton, half of it richly coloured and half bleached stark white

The SPOFs You Did Not Design

Single points of failure are one of the oldest concepts in systems engineering. They are also one of the most misunderstood in modern architectures. Cloud-native platforms were designed to eliminate them. Redundancy, replication, distribution across zones and regions. The assumption is that if no single component is irreplaceable, the system has no SPOF. That assumption is structurally incomplete. What changed is not the presence of single points of failure. What changed is where they live, how they manifest, and why they remain invisible until an incident exposes them. ...

May 7, 2026 · 10 min · 2081 words · Andre Rocha
Thousands of starlings massed into one dark shifting cloud above a twilight marsh, a looser scatter of larger birds flying beneath them

What Breaks When: An Interactive Cluster Failure Explorer

Let me show you something. Architecture diagrams are static. They show components, boundaries, and arrows. What they do not show is what happens when one of those components fails. This one does. I laid out five OpenShift cluster patterns side by side. None of them are hypothetical, I pulled each from real production environments: single cluster with multiple node pools, hosted control planes, ACM-federated fleets, air-gapped stacks, and isolated compliance zones. Click any component and watch what propagates. The red pulse is the blast radius: everything impacted, directly or indirectly, when that component fails (FN-0002). ...

April 28, 2026 · 3 min · 550 words · Andre Rocha
Honeycomb of hexagonal cells sharing common walls, covered with honeybees, the cells holding honey, coloured pollen and capped brood

Cost Optimization vs Risk Concentration in Hosted Control Planes

Hosted control planes are presented as a cost optimization strategy. They are also a risk consolidation strategy. The industry treats these as separate conversations. One belongs to FinOps reports. The other belongs to architecture reviews. ...

April 16, 2026 · 9 min · 1766 words · Andre Rocha
Cutaway of a forest floor: mushrooms above the soil line, a branching white mycelial network spreading through the dark soil beneath

The Hidden Reliability Risks in Multi-Cluster Kubernetes

Multi-cluster Kubernetes is often introduced as a solution to failure. In practice, it does something more subtle. It changes the shape of failure. Failures do not disappear. They stop being local, predictable, and contained. They become distributed, indirect, and delayed. The most dangerous part is not the failure itself. These failure modes share a pattern: they rarely appear in architecture diagrams, do not violate best practices, and only become visible under specific lifecycle events. ...

March 31, 2026 · 8 min · 1607 words · Andre Rocha
An orb-weaver spider at the center of its web, radial threads converging on the hub

Cloud-Native, Same Old Fragility

Modern systems are distributed. But fragility didn’t disappear. It just became harder to see. They run across clusters, regions, providers . They are observable, containerized, orchestrated . ...

March 10, 2026 · 3 min · 547 words · Andre Rocha
Octopus on the seabed, arms extended

Translating OpenShift Health into Business Risk

The gap no one owns Most OpenShift environments can report their health status with precision. Very few can report their risk position with confidence. Clusters expose thousands of signals: node conditions, operator status, etcd latency, certificate countdowns. The data exists. What rarely exists is a structured translation layer between platform health and business risk. ...

February 17, 2026 · 11 min · 2258 words · Andre Rocha
Bird nest woven from grass, moss and lichen, holding a clutch of speckled eggs, with a blue tit perched on the rim

Why Most OpenShift DR Strategies Fail at Executive Level

Most enterprise OpenShift disaster recovery strategies are designed to satisfy audits, not to survive real incidents. They describe recovery procedures, declare RPO and RTO targets, and satisfy audit checklists. What they rarely do is demonstrate recovery capability under realistic conditions. This distinction matters more than it appears. Having a D.R. plan and having D.R. capability are fundamentally different things. The first is a document. The second is a measurable organizational competence that requires investment, testing, and continuous validation. ...

January 27, 2026 · 11 min · 2185 words · Andre Rocha