Sixteen years leading every layer of IT infrastructure. Now I use AI to build and scale reliable systems faster than ever.
Site reliability and infrastructure leader who has built IT, DevOps, and SRE organizations end to end: founded one from nothing, then led platform-wide SRE governance at a global fintech. Owning every layer myself is what makes SRE work: I already know what breaks, and why. None of it is solo. I partner with the business units I serve and lead teams of 3 to 10+ across London, India, Mexico, and both US coasts.
A tool is only as good as the person using it.
AI didn't hand me new judgment, it handed me leverage. Sixteen years running every layer of IT taught me what "reliable" actually means, and what breaks when it isn't. That's what's behind what's next: AI extending SRE judgment, not replacing it.
Purpose built, not off the shelf: custom agents on a real agent development kit (ADK), talking to each other over the agent to agent (A2A) protocol. What started as SRE automation is built to grow into the operating model for the whole tech org: agents propose, a human approves, across security, compliance, IaC, incident command, and CAB review, not just reliability. Today those domains depend on a handful of senior engineers whose judgment doesn't scale past them.
Watches latency, traffic, errors, and saturation, and holds new services to instrumentation standards.
Tracks error budget burn and ties it to release velocity: burn too fast, and risky changes get held.
Flags recurring manual work and proposes the automation that removes it.
Forecasts demand against headroom and flags services that need to scale before it's an incident.
Scores changes against the IaC baseline and incident history, a first pass CAB review before a human signs off.
Stands up the incident channel, assigns roles, and builds a live timeline the moment an incident is confirmed.
Drafts blameless postmortems and tracks action items, feeding lessons back into the orchestrator.
Continuously scans for misconfigurations, exposed secrets, and hardening drift, not just at audit time.
Maps NIST, PCI, ITAR, and MPAA controls to live evidence, so audit prep is a review, not a scramble.
Enforces policy as code across Terraform and ArgoCD, blocking guardrail violations before a CAB review.
Summarized at a high level here. Happy to go deeper live.
Looking for my next SRE or platform leadership role: Director, Head of SRE, or a principal level hybrid, somewhere reliability and automation are the mission. If that's your team, let's talk.