4.5 KiB
On-call by default
- 期号: SRE Weekly Issue #301(2021-12-19)
- 作者: Chris Evans — incident.io
- 链接: https://incident.io/blog/on-call-at-incident-io
简介
These folks put everyone on-call by default, and also pay them extra automatically for each shift and even covering for coworkers.
正文
December 16, 2021 — 3 min read
Like many SaaS businesses, we use our own on-call software to provide 24x7 cover if there are problems with incident.io. We have a 'pager' which will alert the relevant person if something unexpected happens in our app, so that they can investigate and fix it if needed.
Note: This was adapted from an internal document we wrote about how we think about on-call at incident.io.
We're building a product that people depend on 24x7, all year around. It's important it always works, that means we need to support it around the clock. During office hours this is a shared responsibility across the whole team, but to limit the impact out of hours, we have a dedicated person 'holding the pager'.
Being on-call doesn't come without its benefits. Designing our on-call schedules around this principle tightens the feedback loops between shipping and running. This helps us to make pragmatic engineering decisions and provide a healthy tension between shipping new code, and supporting and improving what we have.
Additionally, our product is designed, partly, to support folks who are on-call. There's no better way for us to empathise with our customers than to do the job ourselves — as detailed in being on-call at incident.io.
As an incentive, and to compensate for the inconvenience of having to remain close to your laptop, we'll pay a fixed amount per week to anyone who's on-call.
We'll calculate on-call compensation automatically from our schedules, and take overrides into account too — down to the minute, so if you cover someone for an hour while they go to the shops, you'll be paid for that time.
By compensating on-call we also aim to make overrides feel more fair, and avoid the need for more complex swaps of time. If someone offers to cover a day of your shift, they'll be paid for it so there's no need to feel indebted.
On-call payment is not expected to cover any time you spend working outside of hours. If you're paged and end up working in your evening, you should take time off in lieu. We trust you to manage this time yourself.
Being on-call unavoidably has an impact on your home life, but we want to provide the best possible experience. Here's a few ways we'll collectively help each other:
- Following on-call best practices , if you're paged during the night we'll find replacements for the next day and expect you to take the time back.
- If you've got something going on at home we encourage you to ask for support. As we grow, it's going to become increasingly rare for there to be times when nobody is at home and near a laptop, so just ask!
- We'll always make sure you're supported. You'll always have the three founders as a backstop, and we collectively agree that it's ok to opportunistically page anyone if you need some support.
- We'll make sure our systems are set up to tell you when you need to do something. You shouldn't be watching Slack to keep things going.
Chris Evans
Co-Founder & Field CTO
I'm one of the co-founders, and Field CTO here at incident.io.
Today we're launching Investigations: agentic root cause analysis that starts the moment you're paged, figures out what broke and why, and works with your team through to resolution. Here's what we built, what's powering it, and why it took some time to get right.
Pete Hamilton
August 5, 2026
PagerDuty published a new comparison table about incident.io. Once again, it describes a product we don't recognize. So once again, we're correcting the record, row by row, with receipts.
Tom Wentworth
July 28, 2026
Today, we're launching the Opsgenie Rescue Program to make that landing soft: simplified migration and free overlap so you never pay two vendors at once.
July 9, 2026
Ready for modern incident management? Book a call with one of our experts today.
- All-in-one incident management
- Our unmatched speed of deployment
- Why we’re loved by users and easily adopted
- How we work for the whole organization