Site Reliability Engineer · Production Systems

Keep the signal.
Lose the noise.

An SRE focused on resilient platforms, useful observability, safe delivery and calm incident response. I turn operational risk into systems teams can reason about.

SERVICE
HEALTH
Telemetry SLOs Deployments Incident response
01 / Reliability ethos

Calm systems
scale better.

Reliability is not a heroic act at 2 a.m. It is the compound effect of good defaults, measurable objectives, fast feedback and systems designed for failure.
01 · Measure

SLOs over vibes

Translate user impact into service objectives and error budgets that guide engineering trade-offs.

02 · Observe

Signal over dashboards

Design telemetry around questions responders actually need answered during normal operation and incidents.

03 · Automate

Remove repeat toil

Turn recurring operational work into safe, documented automation with clear ownership and rollback paths.

04 · Learn

Blameless by default

Use incident reviews to improve systems, interfaces and detection rather than hunting for individual fault.

02 / Operations

Production
under control.

Selected operational stories show how I reduce blast radius, shorten recovery and make the next incident easier to handle.

P1 · 14 MIN

Checkout latency spike

Traced a cross-region dependency regression, introduced saturation alerts and redesigned retry behaviour.

MTTR down 41%
P2 · 28 MIN

Queue backlog storm

Added consumer lag SLOs, adaptive capacity and a replay-safe recovery runbook.

TOIL down 6 H/WK
P2 · 19 MIN

Deploy health regression

Built progressive delivery checks around golden signals and automated rollback conditions.

FAILURES down 32%
03 / Expertise

The reliability
toolbox.

Expertise is represented as signal clusters rather than artificial percentage bars.

01 · PLATFORM

Cloud and Kubernetes

Design resilient workloads, capacity boundaries, autoscaling and failure-aware platform patterns.

02 · OBSERVABILITY

Metrics, Logs and Traces

Build telemetry that connects service behaviour to user impact and actionable operational signals.

03 · DELIVERY

Safe Deployments

Progressive delivery, release health, rollback strategy and automation for confident change.

04 · RESILIENCE

Chaos and Failure Modes

Model dependencies, test recovery paths and make failure behaviour explicit before production finds it.

05 · INCIDENTS

Response and Recovery

Lead calm incident coordination, technical diagnosis, communication and durable follow-through.

06 · ENGINEERING

Go, Python and IaC

Automate operational workflows with maintainable code, infrastructure as code and strong defaults.

04 / Selected systems

Reliability
in practice.

Case studies focus on the operating model, failure surface and measurable outcome, not a list of tools.

01 · Platform resilience

Multi-region platform with graceful degradation

Reworked service dependencies, traffic policies and recovery automation for a customer-facing platform serving peak traffic across three regions.

99.95 target availability
37 % less incident impact
02 · Observability

One signal model for 40+ services

Standardised golden signals, ownership metadata and alert quality so on-call engineers could distinguish symptoms from causes.

03 · Delivery

Release guardrails that know when to stop

Progressive rollout checks tied to latency, errors and saturation with automated rollback for unhealthy releases.

05 / Contact

Let us make production
boring.

If you are scaling a platform, improving on-call, defining SLOs or preparing for a reliability review, send a note.

✉ arjun.mehta@example.com
◈ Bengaluru, IN
in GH