Monitoring payloads are rarely born ready for incident response. One system calls a serious event critical, another calls it disaster, and a third sends severity: 5 because apparently numbers needed more emotional range.
Routing has the same problem. A shared Alertmanager may send events for databases, queues, APIs and infrastructure through one integration. The useful ownership facts are already present in labels, but a fixed route cannot always express the whole decision:
Incoming event
-> inspect source, labels and payload
-> normalize fields
-> select team, route and service
-> choose priority and policies
-> set grouping and deduplication
-> continue through the normal alert lifecycle
That decision layer is what Global Orchestration provides in IncidentRelay. It belongs to a group and runs before a service-specific orchestration. It is the right place for rules that decide ownership or apply consistently across several services.
It is not a replacement for routes, services, escalation policies or alert storage. It prepares an event and selects how those existing parts should handle it. Think air-traffic control, not a second airport.
Global and service orchestration solve different problems
The distinction is about what you know when the rule runs.
| Scope | Use it when | Good examples |
|---|---|---|
| Global | Ownership is not known yet, or the rule applies across multiple services. | Select a team from labels, normalize severity, apply a group-wide noise rule. |
| Service | The service is already known and the decision is specific to it. | Choose the database escalation policy, delay one backup warning, queue service diagnostics. |
The runtime order is deliberate:
Integration authentication and normalization
-> group Global Orchestration
-> selected service orchestration, if a service is known
-> normal IncidentRelay lifecycle
-> alert group, child alert, escalation and notifications
A global rule can select a service. IncidentRelay can then run that service's orchestration for the detailed, local decisions. This keeps the global layer from becoming a 900-line constitution for every team in the company.
Global rules should answer “where does this belong?” and “what common shape should it have?” Service rules should answer “how does this service want to handle it?”
The practical use case: one intake, several owners
Suppose one Alertmanager integration receives alerts from the whole production platform. The payload already contains useful labels:
{
"source": "alertmanager",
"title": "High error rate",
"severity": "fatal",
"status": "firing",
"labels": {
"environment": "production",
"application": "checkout",
"component": "database",
"instance": "checkout-db-2",
"customer_impacting": true
}
}
A global definition can turn that generic payload into an operational decision:
- Convert
fatalto the team's canonicalcritical. - Select the Payments team because
application=checkout. - Select the Checkout Database service because
component=database. - Set priority
P1because the event is both critical and customer-impacting. - Build stable group and deduplication keys.
- Stop global sibling rules and continue into the selected service's lifecycle.
The result is not merely “send a message to a channel”. It is a normalized event with explicit ownership, policy context and stable identity before the alert lifecycle starts doing expensive and visible things.
Build rules as a pipeline, not a lottery
Rules run in their displayed order, and earlier actions can change values that later rules inspect. That makes orchestration powerful, but it also means order deserves more thought than “this card looked lonely, so I dragged it upward”.
A maintainable global definition usually has phases:
1. Normalize integration-specific values
2. Add or clean common labels
3. Select ownership: team, route and service
4. Select priority and policies
5. Set group_key and dedup_key
6. Apply narrow suppress, pause or drop decisions
7. Queue optional automation
For example, rule one can translate disaster into critical. Rule two can then test only the canonical severity. Without this separation, every routing rule needs to remember every creative spelling ever invented by every monitoring vendor.
Monitoring A: critical
Monitoring B: disaster
Monitoring C: SEVERE_PROBLEM
Global Orchestration: “You are all critical. Please form one orderly queue.”
Use structured labels for ownership
Prefer stable fields such as labels.environment, labels.application, labels.component and event.source. Avoid routing from human prose unless the integration gives you no better option.
| Condition | Quality | Why |
|---|---|---|
labels.application equals checkout | Good | Explicit, testable and independent of title wording. |
labels.environment equals production | Good | Clear operational boundary. |
event.title contains DB | Fragile | Titles change, abbreviations collide and humans enjoy punctuation. |
raw.alerts.0.labels.namespace | Sometimes useful | Precise, but coupled to one integration's raw payload. |
Raw payload fields are available when needed, but normalized event and labels fields are easier to simulate, review and reuse across integrations.
Choose ownership without creating invalid combinations
Global Orchestration can select a team, route or service, but these objects still need to agree with each other. IncidentRelay validates the combination at runtime.
Typical rejected combinations include:
- a route from another group;
- a Sentry route for an Alertmanager event;
- a service owned by a different team than the selected route;
- an escalation or notification policy belonging to another team;
- an object that was disabled or deleted after the orchestration was published.
In hybrid compatibility mode, an invalid candidate can be rejected while the existing lifecycle continues. In orchestration mode, routing is authoritative and a failure can block processing. That difference is why compatibility mode should be chosen with more care than a dark-theme preference.
Understand the two mode switches
Event Orchestration has a runtime mode and a compatibility mode. They answer different questions.
Runtime mode: does this definition affect real events?
| Mode | Behavior |
|---|---|
disabled | Rules do not run on production events. Drafting, validation and simulation remain available. |
shadow | The published version runs and records candidate decisions, but production behavior is unchanged. |
active | Valid decisions are applied to production processing. |
Compatibility mode: who is authoritative?
| Mode | Behavior | Use |
|---|---|---|
legacy | The existing lifecycle stays authoritative. | Safe upgrade default and simulation. |
hybrid | Orchestration sets explicit values; the existing lifecycle fills what remains. | Recommended first production rollout. |
orchestration | Orchestration owns routing and requires a valid route. | Mature definitions after shadow and hybrid review. |
A published definition is required before selecting shadow or active. Also note that active + legacy does not apply orchestration decisions to production. It sounds active because software naming occasionally likes a small practical joke.
A safe rollout that does not require bravery
Do not replace every route on Friday afternoon. Start with one narrow, measurable decision:
- Create a global orchestration in
disabledruntime mode. - Select
hybridcompatibility mode. - Add one rule for a specific source, application and environment.
- Save and validate the draft.
- Simulate a matching event.
- Simulate a near match that must not match.
- Test missing optional labels and a resolved event.
- Publish the version and switch to
shadow. - Review execution traces and shadow metrics on real traffic.
- Switch to
activeonly when candidate ownership and grouping are consistently correct.
Keep existing routes and policies during the first rollout. Hybrid mode lets orchestration make explicit improvements while the existing lifecycle fills gaps. It is a bridge, and unlike many migration bridges, it comes with railings.
Simulation should test the rule's evil twin
Testing only the event you expect is how broad routing rules develop hobbies. For every important rule, prepare at least these cases:
- the exact matching production event;
- a staging event with otherwise identical labels;
- another application with the same severity;
- an event missing an optional field used in a template;
- a resolved version of the same event;
- an event from another integration source;
- a repeat event that must produce the same deduplication key.
The Simulator shows matched rules, condition traces, action changes, selected entities, final disposition and differences from the active version. A simulation can complete successfully and still make a terrible business decision, so verify the output rather than celebrating a green response code.
Grouping is part of routing correctness
Choosing the correct team is not enough if every repeat creates a new alert or twenty hosts collapse into one indistinguishable blob.
group_key: {{ labels.alertname }}:{{ labels.environment }}
dedup_key: {{ labels.alertname }}:{{ labels.instance }}
window_seconds: 900
This groups related instances of one alert in one environment while preserving a child alert per instance. Use stable values. A timestamp in dedup_key is technically valid and operationally equivalent to feeding the incident list after midnight.
Suppress, pause and drop are not synonyms
| Action | What remains | Use it for |
|---|---|---|
suppress | The alert remains visible; notifications and escalation are suppressed. | Known noise that operators may still need to search or correlate. |
pause | A pending event that activates after a delay unless a resolve arrives. | Transient failures that should persist before paging. |
drop | No alert or group is created. | Events with no operational value, such as an exact test heartbeat. |
Never begin with an empty catch-all condition and one of these actions. Especially not drop. A catch-all drop rule is a very efficient alert-fatigue solution in the same way removing the smoke detector is a very efficient low-battery solution.
Use multiple exact conditions for destructive decisions, validate them, replay historical traffic where available, then observe them in shadow mode.
Webhook actions stay asynchronous and controlled
A matching global rule can queue a reusable webhook action for diagnostics, ticket creation or another automation service. IncidentRelay does not execute the HTTP request inside the alert ingestion request. Delivery is asynchronous and handled by the scheduler with retries and security controls.
Simulation and shadow mode never send webhook actions. This is important: “test safely” should not translate into “open 300 real tickets with TEST in the title”.
What to review in shadow mode
Do not treat a low difference count as automatic approval. Open representative executions and inspect:
- unexpected catch-all matches;
- team, route or service changes;
- route/source mismatches;
- priority and policy changes;
- unstable group or deduplication keys;
- missing optional template fields;
- events that would be suppressed, paused or dropped;
- later rules undoing earlier actions;
- rules that never match at all.
Shadow metrics tell you where candidate behavior differs. Execution traces tell you why. You need both: one finds the suspicious neighborhood, the other tells you which house is on fire.
A maintainable rule set has an owner
Global orchestration affects many services, so treat it as shared production code even though the Builder does not ask you to compile anything.
- Give rules business names such as “Route production checkout databases”.
- Describe why the rule exists and who owns the underlying convention.
- Put specific rules before broad defaults.
- Use
stopafter mutually exclusive routing matches. - Use
continuefor independent normalization and enrichment. - Record a publication comment for meaningful changes.
- Review execution samples after integrations add or rename labels.
- Roll back by publishing a new copy of a known-good historical version.
Published versions are immutable, which gives executions a reliable definition to point to. A rollback creates and publishes a new copy rather than quietly rewriting history. This is less magical and considerably better for incident review.
The launch checklist
- The global scope is used for ownership or genuinely cross-service behavior.
- Conditions use stable normalized fields where possible.
- Selected teams, routes, services and policies belong together.
- Rule order is documented and intentional.
- Group and deduplication keys are stable.
- Matching, near-matching, missing-field and resolved events were simulated.
- Destructive actions have narrow conditions and separate review.
- The version spent enough time in shadow mode to see representative traffic.
- Execution traces show the expected ownership and disposition.
- A known-good version is available for rollback.
What Global Orchestration buys you
The real benefit is not having more rules. It is having one explainable place where inconsistent monitoring facts become consistent operational decisions before paging begins.
A good global layer makes service ownership explicit, reduces duplicated routing logic, normalizes vendor-specific payloads and gives risky changes a simulator, immutable versions, shadow evaluation and traces. The event may still announce that production is on fire, but at least it arrives at the correct fire brigade with a readable address.
For the complete Builder workflow, conditions, actions and examples, read the Event Orchestration user guide. The Event Orchestration API guide covers automation and control-plane endpoints.
Want to make this less theoretical?
IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.