Blog

Notes from the on-call rotation

Field notes on incident response, escalation design and running SkyLogs in production.

Featured · Practice

Alert fatigue is an ownership problem

Reducing noise is only half the work. Until every signal has an owner, quiet dashboards just hide the same risk.

2026-07-28 · 6 min

  • On-call

    Designing escalation chains that actually get answered

    Timeouts, fallbacks and manager steps: how to build a chain your team trusts at 3am.

    2026-07-09 · 8 min

  • Deployment

    Self-hosting SkyLogs on Kubernetes

    A walkthrough of the Helm chart, persistence, and safe upgrade paths.

    2026-06-22 · 11 min

  • Concepts

    Alert vs. incident: why the distinction matters

    An alert is a signal. An incident is a response with an owner, a clock and an end.

    2026-06-03 · 5 min

  • Notifications

    Routing notifications by severity, not by habit

    Phone for critical, Slack for warning, email for info — and why the mapping matters.

    2026-05-19 · 7 min

  • Practice

    Measuring on-call load without punishing people

    Fair rotations start with visible data: pages per shift, night interrupts, handoffs.

    2026-05-02 · 9 min

Prefer to see it working?

The live demo walks through a full incident, from alert to resolution.