DevonRay
IT & Technology - Site Reliability Engineer

Keep the
lattice alive.

I engineer dependable production systems by connecting observability, automation, incident response and resilient architecture into one operational mesh.

99.99%SLO TARGET
EDGE / API
QUEUE / WORKERS
DATA / CACHE
OBSERVABILITY
CONTROL PLANE
01 / Operating profile

Reliability is a connected system.

Production is a living dependency graph: signals reveal weak links, automation reduces toil, and every incident feeds the next improvement.

99.99%target availability
41minutes of toil removed / week
23runbooks hardened
17production services supported
02 / Expertise lattice

Signals, safeguards, recovery.

Capabilities shown as modular nodes rather than generic percentage bars, reflecting how SRE work compounds across the stack.

NODE / 01 📊

Observability

Metrics, logs, traces, SLOs and alert design that turn production behavior into a readable signal.

NODE / 02 🚨

Incident Response

Calm triage, escalation, coordination and blameless learning loops.

NODE / 03 ☁️

Cloud Reliability

Capacity, scaling, health checks and failure-aware architecture.

NODE / 04 ⚙️

Automation & Platform Engineering

Runbooks, CI/CD safeguards, infrastructure-as-code and self-service paths that make the reliable path the easy path.

03 / Experience

From alert to durable fix.

A practical SRE career arc built around production ownership, service health and engineering systems that reduce repeat incidents.

Senior Site Reliability Engineer · Northstar Platform

Own reliability for customer-facing services, define SLOs, improve observability, lead incident reviews and partner with product teams on resilience work.

Production Engineer · Meridian Cloud

Automated infrastructure workflows, introduced deployment safeguards and built recovery exercises for distributed workloads.

Systems Engineer · VectorWorks

Supported Linux fleets, monitoring, release operations and the first generation of infrastructure-as-code practices.

04 / Selected work

Proof under pressure.

Representative reliability initiatives framed around the failure mode, intervention and production outcome.

CASE / 001

Alert topology reset

Rebuilt noisy alert routing around service ownership, SLO burn and dependency context, reducing low-value pages while making real incidents faster to diagnose.

OBSERVABILITYSLOsINCIDENTS
CASE / 002

Recovery rehearsal grid

Created repeatable failure drills for queues, databases and regional dependencies so recovery assumptions could be tested before they mattered.

RESILIENCEDR
CASE / 003

Toil-to-code pipeline

Converted recurring operational checks into automated workflows with guardrails, reducing manual release and maintenance work.

AUTOMATIONPLATFORM
05 / Contact

Make your next incident easier to survive.

For SRE, platform engineering, observability, production readiness and reliability leadership conversations.

✉ devon@example.com
◈ Dublin, Ireland
in GH