Operations

Governing a Multi-Instance Azure SQL Estate

Some of the most expensive operational problems become visible as technical failures, but their causes often begin much earlier.

Weak visibility. Unclear ownership. Missing controls. Decisions made too late.

In my most recent role, I was trusted with operational ownership across a multi-instance Azure SQL estate spanning EMEA, the Americas, APAC and the South Pacific, supporting hundreds of databases.

The technical scale mattered, but the greater responsibility was making the environment understandable enough to govern.

I built and operated a governance dashboard that gave IT leadership clearer visibility of capacity, cost, risk and operational health. That visibility helped us identify capacity risks before they became service-impacting incidents, recover significant operational headroom and surface meaningful opportunities for cloud cost optimisation.

What stayed with me was not simply the technical work. It was the pattern behind it.

Many operational failures are not caused by technology alone. They emerge from the gaps around it.

Gaps in visibility, ownership, controls and decision-making as environments grow in scale and complexity.

That changed how I think about good operations.

What operational excellence actually requires

Operational excellence is not only about how quickly a team responds when something breaks. It is also about whether the environment has been designed and governed well enough to identify deteriorating conditions before customers experience them.

One thing I have learned from operating in global enterprise environments is that reliable cloud and SaaS operations do not happen by accident. They require deliberate operational architecture:

  • Enough visibility to understand what is happening.
  • Clear accountability for what matters.
  • Effective controls around change and risk.
  • The judgment to act before technical conditions become business disruption.

The goal is not to eliminate every incident. No operational environment can promise that.

The goal is to make complex environments easier to understand, safer to change, faster to recover and more predictable as they scale.

That is the standard of operational maturity I believe organisations should be building towards.