From alert to mitigation
From alert to mitigation
An incoming alert starts the response. Connected telemetry and environment context help Aiden explain it; responders use the findings to decide what to do and check whether the action worked.
1. Receive the signal
Alert ingestion supports webhooks and, for supported sources, scheduled or manual imports. Each integration instance has its own delivery settings. For example, production and staging Grafana instances can use different schedules and filters.
Filters answer different questions: what the provider returns, what Aiden ingests, and which ingested alerts should start an investigation automatically. Queue filters only change the list you see.
2. Triage and correlate
On Alerts, compare the provider's Monitor severity with Aiden's Triage priority. Read the incident summary and its explanation. Root signal identifies a suggested starting point in a linked group; Downstream effect points toward a related alert to check first. These are investigation leads to validate against evidence.
Automatic investigation can start from ingestion when enabled. It does not require a responder to first select a summary card or manually categorize the alert.
3. Investigate with context
Discovery gathers environment context from connected integrations. During an investigation, Aiden can consult available telemetry, repository context, and runbooks. A remote runner supplies private-network access for supported tools when configured.
In the investigation workspace, use evidence to check the finding and events to understand what ran. Ask what remains unverified when a datasource, permission, or runner is missing.
4. Review and carry out mitigation
A finding or proposed command is not evidence that a change ran. Review the target service and environment, expected effect, risks, rollback, and required approval. Execution depends on the configured tools, credentials, workflow, and permissions. If the required action is unavailable in SRE, hand the plan to the operator or change workflow that owns it.
Useful follow-up:
Propose the smallest mitigation for this service. State what evidence supports it, what approval is required, how to roll it back, and which checks will confirm recovery.
5. Verify and close
Check the relevant metrics, logs, alert state, and user-facing behavior after the change. Compare against the incident window and a healthy baseline. Record the action and verification in the conversation, then resolve the investigation threads when the response is complete.
Resolving threads records the case status in SRE. It does not itself repair the service or resolve an incident in an external system. Keep external incident records, such as FireHydrant, aligned through your team's incident process.