7 ms·
Launch HN: Rootly (YC S21) – Manage Incidents in Slack
Hi HN, Quentin and JJ here! We are co-founders at Rootly (https://rootly.com/ https://rootly.com/), an incident management platform built on Slack. Rootly helps automate manual admin work during incidents like the creation of Slack channels, Jira tickets, Zoom rooms & more. We also help you get data on your incidents and help automate postmortem creation.
We met at Instacart, where I was the first SRE and JJ was on the product side owning ~20% GMV on the enterprise and last-mile delivery business. As Instacart grew from processing hundreds to millions of orders, we had to scale our infrastructure, teams, and processes to keep up with this growth. Unsurprisingly, this led to our fair share of incidents (e.g. checkout issues, site outages, etc.) and a lot of restless nights while on-call.
This was further compounded by COVID-19 and the first wave of lockdowns. We surged in traffic by 500% overnight as everyone turned to online grocery. This highlighted our need for a better incident management process as it stressed every element of it. Our manual ways of working in Slack, PagerDuty, Datadog, simply weren’t enough. At first, we figured this was an Instacart-specific problem but luckily realized it wasn’t.
A few things here. Our process lacked consistency. Depending on who was responding and their incident experience it varied greatly. Most companies after they declare an incident rely on a buried-away runbook like on Confluence/Google Docs to try and follow a lengthy checklist of steps. This is hard to find, difficult to follow accurately, slow, and stress inducing. Especially after you’ve been woken up to a page at 3 am. We started working on how to automate this.
Fast forward to today, companies like Canva, Grammarly, Bolt, Faire, Productboard, OpenSea, Shell use Rootly for their incident response. We think of ourselves as part of the post-alerting workflow. Tools like PagerDuty, Datadog act like a smoke alarm to alert you to an incident, which hand off to Rootly so we can orchestrate the actual response.
We’ve learned a lot along the way. We realized the majority of our customers use the same 6 (Slack, PagerDuty, Jira, Zoom, Confluence, Google Docs, etc.) tools, follow roughly the same incident response process (create incident → collaborate → write postmortem), but their process varies dramatically. The challenge in changing these processes is hard.
Our focus in the early days was build a hyper opinionated product to help them follow what we believe are the best practices. Now our product direction is focused on configuration and flexibility, how can we plug Rootly into your already existing way of working and automate it. This has helped our larger enterprise customers be successful with their current processes being automated.
Our biggest competition is not PagerDuty/Opsgenie (in fact 98% of our customers use them) or other startups. Its internal tooling companies have built out of necessity, often because tools like Rootly didn’t exist yet. Stripe (https://www.youtube.com/watch?v=fZ8rvMhLyI4 https://www.youtube.com/watch?v=fZ8rvMhLyI4) and GitLab (https://about.gitlab.com/handbook/engineering/infrastructure/incident-management/ https://about.gitlab.com/handbook/engineering/infrastructure...) are good examples of this.
Our journey is just getting started as we learn more each day. Would love to hear any feedback on our product or anything you find frustrating about incident response today.
Leaving you with a quick demo: https://www.loom.com/share/313a8f81f0a046f284629afc3263ebff https://www.loom.com/share/313a8f81f0a046f284629afc3263ebff
- mchusma 4y agoI do t get the pricing on things like this and pingdom. This stuff seems like it should be cheap, like $5/user/mo. But everyone seems to go expensive. There are other industries that are similar, but this always stood out to me as an industry where the pricing never felt right to me.
- kwent 4y agoI appreciate the feedback. We tried to think of our pricing as the value a user would get out of the tool. The amount of time and headache we'd save them when actually responding to an incident. We've found a lot more pushback for smaller sized companies (e.g. <20 eng) but have also realized the challenges of managing incidents are less pronounced. Just curious for my own learning, is the thinking behind "should be cheap" motivated by potentially not everyone would need access to tools in monitoring, alerting, response?
- danpalmer 4y agoFor non-technical purchasers, it makes sense to price by the value the organisation will get out of the tool. However for tools that technical users are involved in you have to fight against the "I could build this myself" factor. There are lots of tools that are basically CRUD apps, or maybe CRUD with a chat interface in this case, which are fundamentally straightforward to build a first-pass version of. I'm sure the product here is far better than a first-pass version, but it's an uphill battle to justify that when the pricing is on the value to the user, rather than based on the cost to build. Another complicating factor for this market in particular is that there are often two types of users: regular and infrequent. In my experience tools like this would be used heavily by the engineering team, but there was value in having everyone in the org have access to the tool. There may be 10x the number of non-engineers, but they're often worth 1/10th or less to have on the platform. Each person isn't worth having by themselves but having everyone there is worth something. Nickel-and-diming customers on the basis of lots of users who rarely use the platform isn't great. Edit: also, don't have a pricing page with no pricing on it.
- 4y ago