API response times halved.
Instrumented diagnosis and targeted changes, not a rewrite.
SRE · Backend · Production systems
I work across backend and infrastructure on high-traffic production systems to diagnose problems, deliver the fixes and leave the team able to operate independently.
Eight years building and running complex systems, with end-to-end ownership from diagnosis to operations.
When to call
When a production issue needs ownership, not another recommendation deck.
Incidents recur, recovery depends on a few people, or reliability is not measured clearly.
The averages look acceptable, but users, tail latency or background jobs say otherwise.
The bill grows faster than usage and the trade-offs are hard to quantify.
Kubernetes, PostgreSQL or cloud changes are blocked by production risk.
Backend, PostgreSQL and infrastructure all look like plausible causes. I measure the system end to end and fix the problem where it actually is.
Observability, SLOs, incident response and production responsibilities remain implicit or scattered.
How I work
The right solution, not the most impressive one.
Establish the signal, the baseline and the constraint before choosing a solution.
Make the necessary changes, one controlled step at a time, and verify the result.
Document, automate and transfer the operating knowledge. The team carries on without me.
Results
Instrumented diagnosis and targeted changes, not a rewrite.
The CI bill was also cut to a third.
Including pg_repack and Row Level Security changes.
Expertise
Work is organised around the problem to solve, not around a tool catalogue.
Observability, SLOs, incident response, profiling and end-to-end latency analysis.
APIs and services in Python or Go, asynchronous processing, PostgreSQL, performance and sensitive production migrations.
Kubernetes, cloud, bare metal, CI/CD, infrastructure as code and cost control.
Frequent environments — Python, Go, PostgreSQL, Kubernetes, AWS, GCP, OVH bare metal, Terraform/Pulumi, Cloudflare, Datadog.
A simple first step
A short, focused engagement to establish what is happening in production and determine where to intervene before making broader changes.
Signals and constraints are measured, not assumed.
The main failure modes and technical risks are made explicit.
Changes are ordered by impact, risk and effort.
What should happen next, including when my involvement is not needed.
No commitment to a longer engagement.
Start with a production diagnosticDescribe the problem in a few lines. I reply within two business days and will tell you plainly whether I am the right person.
Thirty minutes, no commitment. Most useful if you already have a number or symptom to examine.
Talk about your production — 30 min