SRE / production systems / calm under pressure

Keep the lights on.

I design reliable platforms, observable services, and humane incident practices that help engineering teams ship boldly without making production fragile.

SIGNAL → RESPONSE → RECOVERY
01 / signal

Reliability is a living system.

I work where software meets operations: turning noisy telemetry into useful signals, failure modes into learning, and repeated toil into automation.

99.97% SERVICE AVAILABILITY

Measured availability for a customer-facing platform.

42RUNBOOKS AUTOMATED

Recovery paths made executable instead of tribal knowledge.

18MIN INCIDENT MTTR

Improvement from alert to verified recovery.

4CORE SLOs

Availability, latency, freshness and correctness kept visible.

02 / craft

Tools are only useful when the system is understood.

Capabilities spanning observability, cloud infrastructure, incident response, automation and platform engineering.

Observability & Telemetry
Kubernetes & Containers
Cloud Architecture
Terraform & IaC
Incident Command
Capacity & Performance
03 / systems

Selected reliability stories.

Projects framed as operational outcomes rather than generic software demos.

01 · OBSERVABILITY-38% noisy alerts

Northstar Control Plane

Unified service telemetry, SLO dashboards and alert routing across a multi-team platform. Introduced signal ownership and burn-rate alerts so teams could act before customers noticed.

02 · AUTOMATION7 min → 90 sec

Recovery Loom

Codified common incident recovery workflows into safe, audited automation with clear rollback paths.

03 · RESILIENCE+2 regions

Failover Atlas

Designed multi-region recovery exercises, dependency maps and game-day scenarios for critical services.

04 / journey

A career measured in calmer incidents.

A production timeline focused on moments of operational impact.

Senior Site Reliability Engineer · Platform Engineering

Own SLO strategy, observability standards, incident readiness and reliability roadmaps for high-traffic services.

Site Reliability Engineer · Distributed Systems

Built deployment guardrails, automated toil, improved capacity planning and partnered with product teams on reliability goals.

Infrastructure Engineer · Cloud Operations

Established infrastructure-as-code, monitoring foundations and repeatable production environments.

05 / contact

Let's make production boring.

Open to conversations about SRE leadership, platform resilience, observability, incident engineering and operational excellence.

✉ avery@example.com
☎ +1 646 555 0142
◈ New York, US