Site reliability engineering · production systems

Fold chaos
into calm.

I build SRE practices that turn complex production environments into systems teams can observe, operate and improve with confidence, one carefully folded layer at a time.

01 / Operating philosophy

Reliability is
a folded system.

Good SRE work connects technical controls with human decisions. I focus on service-level objectives, meaningful telemetry, safe delivery, capacity awareness and incident learning so production becomes more predictable without becoming rigid.

99.97% target service availability
36production incidents led
17min median recovery
58operational tasks automated
02 / Reliability toolkit

Four folds.
One stable shape.

Capabilities are shown as layered operating disciplines rather than generic percentage bars.

FOLD / 01

Observability

Metrics, logs and traces designed around user impact, useful thresholds and actionable diagnosis.

PrometheusGrafanaOpenTelemetry
FOLD / 02

Cloud reliability

Resilient infrastructure with sensible failure domains, capacity planning and repeatable recovery paths.

KubernetesAWSTerraform
FOLD / 03

Delivery safety

Progressive delivery, automated checks and rollback paths that make shipping safer without slowing teams.

CanaryCI/CDSLO gates
FOLD / 04

Incident response

Clear roles, calm communication, practical runbooks and blameless learning loops for high-pressure moments.

On-callICPostmortems
03 / Experience

Systems I have
kept standing.

Selected roles showing a progression from systems operations into reliability ownership and platform engineering.

Senior SRE · Northstar Cloud

Own reliability for multi-region APIs and internal platforms. Introduced SLOs, error-budget policy, burn-rate alerting and automated rollback paths across 30+ services.

Platform Engineer · Relay Systems

Built Kubernetes platform primitives, deployment guardrails and developer self-service tooling while reducing noisy alerts and improving recovery readiness.

Systems Engineer · Meridian Labs

Operated Linux fleets and distributed services, developed capacity models and established the first structured incident-review practice.

04 / Selected systems

Proof in
production.

Case studies framed around the operational problem, the intervention and the reliability outcome.

CASE 02 / MIGRATION

Zero-downtime migration path

Progressive schema changes, canaries and rollback checkpoints for a high-volume service.

CASE 03 / INCIDENTS

Incident command kit

Role cards, automation hooks and communication patterns that reduce cognitive load during outages.

05 / Method

Observe → fold →
test → learn.

A practical reliability loop: understand the signal, make the boundary explicit, automate repeatable decisions, then learn from what production teaches you.

01 / OBSERVE

Find the signal.

Start with user impact, telemetry and failure evidence rather than assumptions.

02 / FOLD

Set the boundary.

Define SLOs, dependencies, capacity limits and explicit failure modes.

03 / TEST

Prove the path.

Exercise rollbacks, recovery procedures and automation before production needs them.

04 / LEARN

Improve the shape.

Use incidents and near misses to make the system safer and the next response calmer.

06 / Contact

Let's make
production calmer.

For SRE leadership, platform reliability, incident readiness, observability or infrastructure work, send a short note.

✉ noah@example.com
☎ +41 44 555 0119
◈ Zurich, CH