The pager fires twelve times. All twelve messages say the database is unreachable, but each comes from a different application instance. Grouping them sounds sensible. Then a disk fills on one host, another region loses connectivity and an alert recovers while the rest are still firing. Suddenly “put everything together” is a surprisingly ambitious incident strategy.
The goal is fewer unnecessary interruptions with the same ability to detect, own and resolve independent failures. A quieter phone is useful evidence only when the underlying coverage still works.
Separate identity, grouping and notification policy
| Mechanism | Question it answers | What it must preserve |
|---|---|---|
| Deduplication | Is this an update to the same active alert instance? | The identity and current state of that instance. |
| Grouping | Which distinct alert instances belong in one response context? | The individual children, affected resources and recovery states. |
| Notification timing | When should someone be interrupted about that context? | Urgent changes and an observable delivery path. |
| Silence or inhibition | Which notifications are intentionally suppressed? | The scope, reason and conditions for ending suppression. |
Prometheus Alertmanager documents grouping, inhibition and silences as separate mechanisms. Keep that distinction when designing the rest of the pipeline: combining related messages and deliberately withholding a notification are different decisions.
In IncidentRelay, a dedup_key identifies a concrete child alert, while a group_key determines its group. Two hosts can have different child identities but share a response group. Responders acknowledge and resolve the group; the children retain the individual signals needed for investigation.
Design an identity that survives a retry
Start by writing down what can fail independently. For a filesystem alert, “host” is usually too broad: /var can recover while /data remains full. For a queue, the identity may need the cluster and queue name. For a regional availability check, it may need the affected region.
A conceptual identity for a filesystem condition might contain these stable dimensions:
monitoring source + environment + cluster + alert rule + host + mountpoint
This is a design checklist, not an IncidentRelay key format to paste into a configuration field. Use the integration's documented fingerprint or key mechanism. If you build a custom adapter, encode fields unambiguously, preserve case where it is meaningful and define how missing values are handled. Concatenating arbitrary strings without escaping is an excellent way to invent accidental identity theft for disks.
| Good identity material | Usually poor identity material | Why |
|---|---|---|
| Stable rule ID and resource identity | Rendered summary text | Text changes should not create a new problem. |
| Environment, cluster and resource scope | Only a short hostname reused across environments | Two independent systems can otherwise collide. |
| Mountpoint, queue or other independently recovering object | Current metric value | Values change during the same failure. |
| The same source fingerprint for firing and recovery | Webhook delivery ID, retry number or receipt timestamp | Transport retries are not new alert identities. |
Keep severity out of a custom instance identity when it represents a changing assessment of the same condition. Some monitoring systems derive fingerprints from labels, so changing a severity label can change their identity. Test the actual source behavior instead of assuming your preferred identity model wins.
Do not deduplicate different resources simply because they share a title. “Database unavailable” in two independent clusters is two pieces of work, even when both messages were written by the same optimistic template.
Choose group boundaries from the response
A useful group normally shares an owner, a meaningful failure scope and a next action. Ask whether one responder can investigate and coordinate the children together. If the answer depends on hiding which tenant, environment or region failed, the group is too broad.
Consider a production API in two regions. Ten instances in one region lose their connection to the same database. Grouping those ten symptoms can help. Folding in an unrelated failure in the other region can hide a separate impact boundary and recovery path.
An illustrative IncidentRelay route group_by list is:
["alertname", "severity", "environment", "cluster", "region"]
This combines instances only when those fields agree. Include mountpoint for filesystem conditions when the affected filesystem changes the investigation. Include instance when each host needs separate ownership or remediation. These are starting points for a route with known labels, not a universal recipe.
Missing labels deserve their own test. Adding region to a grouping policy does not create region data in a payload that never had it. Define where that label is supplied, reject or visibly flag incomplete input in your ingestion process, and inspect the resulting keys.
IncidentRelay 2.2.0 defaults to group_by = ["alertname", "severity"] when a route does not specify it. Normal group construction also includes source, team, route and service scope. That is useful separation, but environment and region still need attention when one route and service contain several of them. An orchestration override can change the effective group key; inspect the resulting decision as well as the route form.
Dashboard: “Only one incident.”
Incident title: “Production.”
A worked example: two disks, one response group
Suppose a route groups by alert name, severity, environment, cluster and mountpoint. Two hosts report that /var is full. Their abbreviated Alertmanager alert items look like this:
{
"status": "firing",
"labels": {
"alertname": "DiskFull", "severity": "critical",
"environment": "production", "cluster": "payments-eu",
"instance": "host-a", "mountpoint": "/var"
},
"fingerprint": "payments-eu-host-a-var"
}
{
"status": "firing",
"labels": {
"alertname": "DiskFull", "severity": "critical",
"environment": "production", "cluster": "payments-eu",
"instance": "host-b", "mountpoint": "/var"
},
"fingerprint": "payments-eu-host-b-var"
}
These are two separate objects inside the webhook's alerts array, not a complete HTTP request. The fingerprints above are readable example identifiers; a real sender supplies its documented fingerprints.
| Incoming event | Expected result | Check |
|---|---|---|
| Host A fires | One child in a firing group. | A is visible with its own fingerprint. |
| The same active alert is delivered again | The existing child is updated. | Its ID is unchanged; no second child appears. |
| Host B fires with its own fingerprint | Two distinct children share the group. | Both hosts remain visible. |
| A resolves with the original fingerprint | A resolves; B remains firing. | The group stays open. |
| B resolves | No child remains active. | The group resolves. |
| A fires again after resolution | A new occurrence is created. | History and first-seen time distinguish the recurrence. |
The last row matters. In IncidentRelay 2.2.0, resolved child alerts are final occurrences: a later firing with the same key creates a new child instead of rewriting the resolved one. If the old group is fully resolved, normal grouping creates a new group too. If other children still keep a matching group open, the new occurrence can join that open group.
Recovery is per instance, not per envelope
A webhook batch can contain both firing and resolved children. Process and verify their individual states. One green child is not evidence that every affected resource recovered.
In IncidentRelay, incoming recovery resolves the matching child, and the group resolves when all its children are resolved. An orphan recovery that matches no active child does not create a new group. Manual resolution is different: resolving the group resolves all of its children. Use that action with an explicit understanding of the monitoring state.
For Alertmanager, enable send_resolved: true on the receiver. For Grafana, keep resolved messages enabled. Preserve the identity and routing context between firing and recovery. A recovery event that loses the fingerprint is not a reliable way to close the alert you meant.
Also test a duplicate recovery, delayed recovery, simultaneous retry and late firing event. Stable keys do not by themselves guarantee ordering or exactly-once processing. Check what your specific integration does with timestamps and stale events; do not assume that a resolved incident is immune to every old payload still waiting in a queue.
ACK should survive repetition, but expose a changed problem
Acknowledgement means someone owns the response. It should not become an indefinite promise that every future symptom is already understood.
In the normal IncidentRelay lifecycle, a new firing child can reopen an acknowledged group. Version 2.2.0 has a specific exception: a route grouped only by a non-empty incident_key or labels.incident_key can preserve ACK for a same-key child that does not raise effective incident priority. A priority increase reopens the group and starts a fresh escalation cycle. Additional grouping fields or an orchestration group-key override do not qualify for that exception.
Use a shared incident key only when an upstream process has established the response relationship. A permanent key such as all-production can quietly turn “same incident” into “same company”. Test a genuinely new failure and a priority increase while the original group is acknowledged.
Account for every waiting room in the pipeline
Batching buys time for related alerts to arrive. It also spends part of your detection-to-page budget. Draw the actual path:
rule evaluation and pending period
-> source notification grouping
-> webhook transit and retries
-> IncidentRelay group wait
-> delivery provider
-> responder acknowledgement
For example, a 30-second source wait followed by a 30-second downstream wait can add roughly a minute before provider delivery, even when both systems are healthy. Measure the full path with timestamps; timers may overlap or behave differently on updates.
Alertmanager distinguishes group_wait, group_interval and repeat_interval; see its routing configuration reference. They control initial batching, updates and repeat notifications, rather than redefining alert identity.
IncidentRelay documents its own initial wait and update interval under ALERT_GROUP_WAIT_SECONDS and ALERT_GROUP_INTERVAL_SECONDS. If a group recovers before its first notification, pending firing notification work is cleared. That can remove transient interruptions, so include short-lived but important failures in your acceptance tests. Do not copy timing values between products without checking their semantics.
Related symptoms still need independent visibility
Two alerts can be related without belonging to one group. A database problem and a checkout error-rate alert may have different owners, impact and recovery conditions. Keeping them separate while linking the investigation can be the clearest model.
IncidentRelay's dependency-aware correlation provides possible root-cause and downstream-impact context. It does not automatically merge groups or suppress their notifications. Treat those relationships as investigation clues, not proof of causality. A second independent failure can happen during the first incident; production has never signed an exclusivity agreement.
If you deliberately suppress downstream notifications elsewhere, define how that suppression ends and how an independent downstream failure remains detectable. Recovery of the supposed root cause is a useful moment to check every remaining symptom.
Prove the policy with a small acceptance matrix
Use an isolated test route and a non-paging destination. Capture the current configuration, send controlled payloads and compare the actual child IDs, group IDs, states and deliveries with explicit expectations.
| Test | Acceptance condition |
|---|---|
| Retry the same active payload, including concurrent delivery. | One active child represents the instance; unexpected duplicate children or pages are investigated. |
| Change only the metric value or description. | The update retains identity when those fields are not part of the source fingerprint. |
| Change the resource, environment or failure boundary. | The independent condition remains a separate child and, where the policy requires it, a separate group. |
| Remove a required grouping label. | The missing scope is observable and does not silently combine unrelated failures. |
| Resolve one of two active children. | The other child and its group remain active. |
| Resolve every child, then fire again. | A new occurrence is visible without erasing the closed history. |
| Add a new child or increase priority after ACK. | Ownership and escalation match the documented lifecycle, including any incident-key exception. |
| Deliver recovery out of order or replay it. | The observed behavior is understood; stale-event handling has an explicit owner if it needs improvement. |
| Trigger a short critical failure and an unrelated failure during an existing incident. | The required notification arrives within the team's measured response budget. |
Save the payloads and expected outcomes beside the route's operating notes. Rerun them after changing labels, fingerprints, grouping rules, adapters or notification timing. A screenshot of one successful test message is not a lifecycle test.
Measure what became quieter
Track at least four units separately: incoming deliveries, stored alert occurrences, groups and human notifications. Deduplication can update an existing row, so the number of stored child alerts is not necessarily the number of webhook deliveries. Use sender or intake telemetry for transport volume instead of treating a database row count as an event counter.
For each rollout, compare a similar observation window and examine:
- pages per incident and time from first symptom to first useful page;
- groups containing unrelated resources, owners or recovery paths;
- critical children added to acknowledged groups and whether responders noticed;
- groups still open after expected recovery, and occurrences immediately following resolution;
- missing fingerprints or grouping labels, failed webhook deliveries and growing retry queues.
A reduction from forty pages to four is promising only if the four still expose every independently actionable problem. Set rollback conditions before changing production: a missed critical test, an unexpected cross-environment group or a response delay beyond the agreed budget is a reason to stop and investigate.
Roll out one response boundary at a time
- Inventory: choose one noisy route and record its current labels, identity, grouping and delivery timing.
- Design: write down what should update one child, what should create another child and what requires another group.
- Replay: run the acceptance matrix in isolation, including partial recovery and a second independent failure.
- Pilot: apply the smallest change, keep the previous configuration and observe a representative operating period.
- Review: inspect individual groups with responders before extending the policy to another route.
Changing keys mid-incident can affect how later events find their existing state. Plan the transition, verify recovery for already-open alerts and avoid treating a configuration rollback as automatic repair of historical associations.
The useful outcome is a response queue whose groups have clear owners and meaningful boundaries, whose children retain their individual state, and whose notifications tell people when the situation actually changes. Quiet is welcome. Missing evidence is not part of the discount.
Implementation references and next steps
The IncidentRelay behavior above was checked against the published v2.2.0 source, particularly the alert lifecycle and active-instance lookup. Check your installed version before applying a lifecycle assumption.
- Alerts and alert groups: group-by fields, lifecycle and notification settings.
- Alertmanager integration and Grafana integration: fingerprints, payloads and recovery delivery.
- Dependency-aware alert correlation: related groups without automatic suppression.
- Explain Trace: inspect the routing and grouping decisions behind an alert.
- On-call metrics and alert-fatigue reduction: review the operational result after the configuration change.
Want to make this less theoretical?
IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.