SLOs over vibes
Translate user impact into service objectives and error budgets that guide engineering trade-offs.
An SRE focused on resilient platforms, useful observability, safe delivery and calm incident response. I turn operational risk into systems teams can reason about.
Translate user impact into service objectives and error budgets that guide engineering trade-offs.
Design telemetry around questions responders actually need answered during normal operation and incidents.
Turn recurring operational work into safe, documented automation with clear ownership and rollback paths.
Use incident reviews to improve systems, interfaces and detection rather than hunting for individual fault.
Selected operational stories show how I reduce blast radius, shorten recovery and make the next incident easier to handle.
Traced a cross-region dependency regression, introduced saturation alerts and redesigned retry behaviour.
Added consumer lag SLOs, adaptive capacity and a replay-safe recovery runbook.
Built progressive delivery checks around golden signals and automated rollback conditions.
Expertise is represented as signal clusters rather than artificial percentage bars.
Design resilient workloads, capacity boundaries, autoscaling and failure-aware platform patterns.
Build telemetry that connects service behaviour to user impact and actionable operational signals.
Progressive delivery, release health, rollback strategy and automation for confident change.
Model dependencies, test recovery paths and make failure behaviour explicit before production finds it.
Lead calm incident coordination, technical diagnosis, communication and durable follow-through.
Automate operational workflows with maintainable code, infrastructure as code and strong defaults.
Case studies focus on the operating model, failure surface and measurable outcome, not a list of tools.
Reworked service dependencies, traffic policies and recovery automation for a customer-facing platform serving peak traffic across three regions.
Standardised golden signals, ownership metadata and alert quality so on-call engineers could distinguish symptoms from causes.
Progressive rollout checks tied to latency, errors and saturation with automated rollback for unhealthy releases.