Skip to main content
3 min read

From alert to mitigation

From alert to mitigation

An incoming alert starts the response. Connected telemetry and environment context help Aiden explain it; responders use the findings to decide what to do and check whether the action worked.

1. Receive the signal​

Alert ingestion supports webhooks and, for supported sources, scheduled or manual imports. Each integration instance has its own delivery settings. For example, production and staging Grafana instances can use different schedules and filters.

Filters answer different questions: what the provider returns, what Aiden ingests, and which ingested alerts should start an investigation automatically. Queue filters only change the list you see.

2. Triage and correlate​

On Alerts, compare the provider's Monitor severity with Aiden's Triage priority. Read the incident summary and its explanation. Root signal identifies a suggested starting point in a linked group; Downstream effect points toward a related alert to check first. These are investigation leads to validate against evidence.

Automatic investigation can start from ingestion when enabled. It does not require a responder to first select a summary card or manually categorize the alert.

3. Investigate with context​

Discovery gathers environment context from connected integrations. During an investigation, Aiden can consult available telemetry, repository context, and runbooks. A remote runner supplies private-network access for supported tools when configured.

In the investigation workspace, use evidence to check the finding and events to understand what ran. Ask what remains unverified when a datasource, permission, or runner is missing.

4. Review and carry out mitigation​

A finding or proposed command is not evidence that a change ran. Review the target service and environment, expected effect, risks, rollback, and required approval. Execution depends on the configured tools, credentials, workflow, and permissions. If the required action is unavailable in SRE, hand the plan to the operator or change workflow that owns it.

Useful follow-up:

Propose the smallest mitigation for this service. State what evidence supports it, what approval is required, how to roll it back, and which checks will confirm recovery.

5. Verify and close​

Check the relevant metrics, logs, alert state, and user-facing behavior after the change. Compare against the incident window and a healthy baseline. Record the action and verification in the conversation, then resolve the investigation threads when the response is complete.

Resolving threads records the case status in SRE. It does not itself repair the service or resolve an incident in an external system. Keep external incident records, such as FireHydrant, aligned through your team's incident process.