Site Reliability Engineer · Availability by design

Keep the signal alive.

I build resilient platforms, useful observability, and calm operational systems so teams can move fast without making reliability a gamble.

Numbers with operational context.

Metrics focused on service health, recovery, automation, and sustainable operations.

99.98service availability
68% faster incident triage
42% fewer repeat incidents
11minute median recovery

Guardrails over heroics.

Reliability is a product capability: explicit service objectives, observable failure modes, tested recovery paths, and automation that gives engineers room to think.

Design for graceful failure

Make degradation visible, bounded, and easier to recover from than a surprise cascade.

Measure what users experience

Use SLIs and SLOs that connect platform behavior to actual service outcomes.

Automate recurring judgment

Turn predictable operational work into safe workflows and self-healing routines.

Learn without blame

Use incident reviews to improve systems, decision-making, and organizational memory.

Observe the whole path.

Reliability work spans the spaces between components: nodes, connections, health signals, and bounded surfaces.

edge / probes
service / mesh
metrics / traces
storage / recovery

From pager noise to signal.

Project stories that foreground resilience, visibility, recovery, and safer delivery.

Reliability program

Error budgets as a delivery control plane

Introduced service objectives and release guardrails that made reliability trade-offs visible to product and engineering teams.

SLOerror budgetrelease safety
Observability

Trace-first incident response

Unified metrics, logs, and traces around critical user journeys to reduce investigation time.

Resilience

Recovery rehearsals

Turned disaster recovery from documentation into repeatable, measured exercises.

Depth where uptime depends on it.

Capability cards that emphasize operational fluency over arbitrary percentages.

Observability
DEEP PRACTICE
Incident Command
DEEP PRACTICE
Kubernetes & Platforms
STRONG
Infrastructure Automation
STRONG
Capacity & Performance
PRACTICAL
Reliability Strategy
DEEP PRACTICE

A record of reducing uncertainty.

An operational timeline rather than a CV-style dump.

2023 - PRESENT

Senior Site Reliability Engineer · Platform Systems

Own reliability strategy for customer-facing services, improve observability standards, and partner on measurable service objectives.

2020 - 2023

Site Reliability Engineer · Distributed Infrastructure

Built incident workflows, automated common recovery actions, and strengthened deployment and capacity guardrails.

2017 - 2020

Infrastructure & Automation Engineer · Cloud Operations

Focused on repeatable infrastructure, production monitoring, and the operational foundations that enabled SRE practices.

Build a calmer production story.

Available for reliability engineering, observability, platform resilience, and operational excellence conversations.

✉ jonas@example.com
☎ +41 44 555 0163
◈ Zurich, CH