The first ten minutes of an incident,
run the same way every time
The first ten minutes of an incident are usually the same checklist every time: restart the service, pull recent logs, check the last deploy, notify the right channel. Whoever is on call ends up doing it manually under pressure at 3am. We build an agent that runs your documented runbook the moment an alert fires, so a human arrives to diagnostics already gathered, instead of a blank terminal.
Why knowing the runbook once is not enough
Most teams with any operational maturity have written runbooks: documented steps for the incidents that happen often enough to be predictable. A service that occasionally needs a restart. A queue that backs up under a specific condition. A database connection pool that exhausts under a known pattern. The runbook is not the problem. Executing it well under pressure is: at an inconvenient hour, while still half awake, even an experienced on-call engineer is inconsistent. Someone newer to the rotation is even less consistent, having read the runbook once in onboarding and now seeing the real incident for the first time.
Gathering context eats real time too, before any actual fix starts. That means pulling the recent logs, checking what deployed in the last few hours, and confirming which service is actually affected versus which one is just downstream noise. None of that is hard. But doing it by hand takes minutes that matter when customers are affected, and it is the same handful of steps almost every time.
The incident record itself gets inconsistent too. What actually happened, in what order, gets reconstructed from memory afterward for the postmortem. That account is usually incomplete, or slightly wrong, by the time it is written down.
What runs the moment an alert fires
Your documented runbooks get converted into steps the agent can actually execute: which logs to pull, which metrics to check, which deploys to look at. They also cover which specific, pre-approved actions are safe to run automatically. That might be restarting a known-flaky service, clearing a stuck queue, or scaling a resource that is under load. The moment a matching alert fires, the agent gathers that diagnostic context immediately. It runs whatever safe first-response steps your team has explicitly approved, creates the incident channel, and populates it with the gathered context, before a human even joins.
For anything beyond the pre-approved safe actions, the agent stops at a clear handoff point. It pages a person with everything gathered so far. That covers what the symptoms are, what changed recently, what the runbook says to try next, and what it has already ruled out. If the incident does not match any documented pattern, it says so plainly, rather than guessing. It pages a human immediately, and flags the gap for your team to turn into a new runbook afterward. Every action taken is logged in order with a timestamp, building an accurate timeline ready for the postmortem, instead of one reconstructed from memory.
Typical integrations: PagerDuty or Opsgenie for alerting, Slack or Telegram for the incident channel, and your existing logging and monitoring stack for diagnostics.
What stays an engineer’s call
Deciding the actual fix for anything beyond the narrow, pre-approved safe actions stays with the engineer who takes the handoff. The agent’s job is making sure they start that work with full context, instead of losing the first ten minutes to gathering it. Which actions get automatic execution rights is a decision your team makes explicitly, runbook by runbook, before launch. The agent never decides that for itself mid-incident. The postmortem itself, what to actually change afterward, is a team discussion the agent’s timeline supports but does not replace.
What’s scoped and logged
Every action the agent takes during an incident is logged in sequence with a timestamp, building a precise record for the postmortem. Automatic execution is scoped to a specific, pre-approved list of safe, reversible actions. Anything outside that list always goes to a human, with no path for the agent to improvise beyond its approved scope. A kill switch disables automatic action execution in one message, while diagnostic gathering and paging keep running. That matters if a specific automated step is suspected of causing problems.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Single automation | from $1,200 | A small set of your most common incident types, diagnostic gathering, safe first-response actions | 1 to 3 weeks |
| Department package | from $2,800 | Incident runbook execution plus on-call summary reports and exception triage | 3 to 5 weeks |
Running cost is usually $20 to $60 a month in model usage depending on incident frequency.
Related
This pairs well with on-call summary reports to close the loop after the incident is resolved. Exception triage and assignment covers the issues that led up to it. For the public-facing side of a live incident, see status page and incident updates.
Full package details are on the AI agents service page and the automation-everything overview. For how we run incident response on infrastructure we operate ourselves, see the secure infrastructure case study and the factory ERP recovery case study.
Tired of every incident starting with the same ten minutes of manual digging? Get in touch and we will turn your runbooks into something an agent can run.
Tired of doing this by hand? We can take the whole routine off your team, not only this step: Routine takeover, from $400 →
FAQ
Does the agent actually fix incidents on its own?
For a narrow set of safe, reversible, pre-approved actions, like restarting a known-flaky service or clearing a stuck queue, yes. Anything beyond that, it gathers diagnostics, runs the safe first steps, and hands off to a person with full context. It never improvises a fix.
How much does incident runbook automation cost?
From $1,200 for a small set of your most common incident types, live in 1 to 3 weeks. A fuller runbook library across multiple services usually runs $2,000 to $3,000.
What if the incident does not match any documented runbook?
The agent says so explicitly, and gathers whatever general diagnostics it can. It pages a human immediately, and flags the gap so your team can document a runbook for that case afterward.
Which alerting and incident tools does this connect to?
PagerDuty, Opsgenie, or a Slack or Telegram-based on-call setup, connected to your logging and monitoring stack, whatever that already is.
Who decides what counts as a safe automatic action?
Your team, explicitly, runbook step by runbook step, before anything is automated. Nothing gets automatic execution rights without your engineers reviewing and approving that specific step first.