On-call dashboards tend to begin with good intentions and end with a large green number nobody trusts. The usual candidates are MTTA and MTTR: easy to calculate, easy to put on a slide and extremely easy to misread without context.
If you measure only speed, responders learn to acknowledge first and understand later. If you measure only alert volume, a well-grouped incident can look quieter than a harmless event storm. If you rank engineers, the dashboard stops measuring the system and starts measuring how people protect themselves from the dashboard.
A useful scorecard needs three views at the same time:
Workload: How much operational demand arrived?
Response: How quickly did ownership and recovery happen?
Quality: Did the response reduce risk without hiding noise?
None of these is a verdict on an individual. Together they are feedback about alerting, routing, staffing and service design.
Rule zero: decide what one alert means
Before counting anything, separate three units that are often mixed together:
| Unit | What it represents | Question it answers |
|---|---|---|
| Raw alert | One event or alert instance received from monitoring | How noisy is the source? |
| Alert group | Related events grouped into one operational problem | How many distinct problems did the team handle? |
| Page or delivery | A notification sent to a human or channel | How many interruptions did the response process create? |
Five hundred raw events can become ten groups and three pages. Reporting “500 incidents” would exaggerate workload; reporting “three alerts” would hide source noise. Keep the nouns attached to every number. Metrics without units are just confident-looking decoration.
IncidentRelay Service Analytics deliberately uses both layers: AlertGroup for grouped operational metrics and raw Alert records for volume and noise analysis.
Use a scorecard, not one heroic KPI
A small team does not need 47 widgets. Start with a balanced set that catches obvious failure modes:
| Dimension | Primary metric | Guardrail |
|---|---|---|
| Demand | Raw alerts and alert groups per service | Top alert names and after-hours pages |
| Grouping | Deduplication ratio | Check for groups that hide unrelated instances |
| Ownership | MTTA p50 and p95 | Escalation rate and missing acknowledgements |
| Recovery | Time to resolution p50 and p95 | Recurrence and premature resolution |
| Backlog | Open and critical-open groups | Age of the oldest open problem |
| Human load | Pages per shift and after-hours interruptions | Distribution across responders |
IncidentRelay exposes the service-level demand, grouping, response and backlog pieces. Page timing, recurrence and per-responder fairness may require notification exports or a small companion report depending on your deployment. Do not pretend an unavailable metric is zero. Unknown and excellent are very different states.
Alert volume: count problems and noise separately
Raw alert volume tells you how hard monitoring is talking. Alert-group volume tells you how many operational problems reached the response workflow. Plot both for the same period.
raw alert volume = all received Alert records
grouped workload = distinct AlertGroup records
dedup ratio = raw alerts / alert groups
A deduplication ratio of 20 means twenty raw alerts became one group on average. That can be healthy: repeated host instances are correctly attached to one incident. It can also hide a monitor firing every few seconds. The ratio is a clue, not a target.
Review the top alert names beside the ratio. A rapidly rising raw count with stable group count usually points to repeated events, unstable instances or an overly chatty resend interval. A rising group count with flat raw volume may point to broken deduplication keys or excessive dimensions in the grouping key.
Dashboard: “One alert group.”
Raw event counter: “I have seen things you people would not believe.”
MTTA measures ownership speed, not problem understanding
In IncidentRelay service analytics, MTTA is calculated from first_seen_at to acknowledged_at when both timestamps exist. It answers: how long did an alert group wait before a responder acknowledged ownership?
Use the distribution, not only the average:
- p50 describes the ordinary response;
- p95 exposes the slow tail where routing gaps, dead phones and awkward handoffs live;
- average is useful for trending but is easily pulled around by a few extreme incidents;
- coverage tells you how many groups actually had an acknowledgement timestamp and therefore contributed to the metric.
A team with a two-minute p50 and a forty-minute p95 does not have a two-minute response process. It has a fast normal path and a dangerous exception path. Investigate the tail by time of day, service, route, severity and escalation outcome.
Acknowledgement should mean “a human has accepted ownership”, not “a bot clicked the button to improve the chart”. Automatic ACK can produce a magnificent MTTA and a completely unattended outage. Congratulations to the graph; condolences to production.
MTTR needs an explicit definition
IncidentRelay currently calculates the displayed service MTTR values from first_seen_at to resolved_at. That is precisely measurable, but the letters “MTTR” are overloaded across the industry: repair, recovery, resolution and sometimes “time until the ticket looked quiet”.
In reports, label the timestamps instead of relying on the acronym:
Time to acknowledge = acknowledged_at - first_seen_at
Time to alert resolution = resolved_at - first_seen_at
Alert resolution is not always full customer recovery. A monitor may recover before a backlog drains, or an operator may resolve a group after moving the investigation elsewhere. Pair time to resolution with service impact, incident review notes and recurrence. If the number improves while repeat alerts increase, you have optimized the closing ceremony rather than reliability.
Track the tail and the missing data
For both acknowledgement and resolution, put p50 and p95 beside the count of eligible groups. A percentile calculated from four acknowledged groups is not a stable service objective, even if the chart uses a very serious font.
Use a minimum sample note such as:
Checkout API, last 30 days
Alert groups: 42
Groups with ACK timestamp: 36 (86%)
MTTA p50: 2m 10s
MTTA p95: 11m 40s
Resolved groups: 38
Resolution p50: 24m
Resolution p95: 3h 12m
The missing 14% is a metric of its own. It may contain auto-resolved events, alerts handled without acknowledgement or incomplete historical data. Explain it before using MTTA to make staffing decisions.
Open backlog is a pressure gauge
Open groups and critical-open groups show current operational debt inside the selected service set. Count them, but also inspect age and state. Ten fresh warnings after a deploy are different from one critical group that has been firing since last Tuesday.
A practical backlog review asks:
- Which critical group is oldest?
- Which open groups are acknowledged but have no next checkpoint?
- Which groups are waiting on another team or upstream service?
- Which alerts should have resolved automatically but did not?
- Which open items are known non-actionable noise?
Do not force the count down by silencing or resolving items without fixing ownership. A clean queue created by hiding work is a tidy-looking incident.
Current impact and historical analytics are different clocks
IncidentRelay Service Analytics is period-based for alert groups, raw events and response metrics. Its Impact widget shows the service's current computed state and blast radius. It is not a historical impact trend for the selected analytics period.
This distinction matters when reviewing last month. A service can show healthy current impact beside a terrible thirty-day alert history, because it recovered. That is not contradictory. One answers “what is affected now?” and the other answers “what happened during the window?”
Use current impact to prioritize today's open work. Use period metrics to choose systemic improvements. Mixing the two is how last month's outage gets declared healthy because the dashboard is green this morning.
Maintenance suppression should have a numerator
IncidentRelay reports maintenance windows and suppressed alert groups per service. Suppression is useful when it removes expected, non-actionable notifications during controlled work. A high number is not automatically good or bad.
Review it beside total alert groups and change activity:
maintenance suppression rate = suppressed groups / all groups in scope
A rising rate during a migration may be expected. A permanently high rate may mean monitors are poorly aligned with maintenance windows, changes are too noisy, or maintenance has become a socially acceptable permanent silence.
Fairness needs human metrics too
Service analytics explains system demand. It does not by itself prove that the on-call burden is fairly distributed. Add a small shift-level report with:
- scheduled primary and secondary hours;
- pages delivered per shift;
- after-hours interruptions;
- escalations received and missed acknowledgements;
- time spent on major incidents;
- recovery time or compensatory time taken after night incidents.
Review this over eight to twelve weeks, not one unlucky weekend. Do not publish an individual leaderboard. The goal is to rebalance rotations, train more responders and fix noisy services, not to discover who can endure the worst sleep while smiling in a chart.
Set targets by service class
A P1 checkout outage and a development-environment warning should not share the same response target. Define targets by severity, environment, service tier and customer impact.
| Class | Example ownership target | Review focus |
|---|---|---|
| Production P1, tier 1 | Fast ACK with tested escalation | p95, missed ACKs, customer impact |
| Production P2 | Predictable ownership within team policy | p50/p95, backlog age, recurrence |
| Non-production actionable | Business-hours ownership | Volume, routing and whether paging is necessary |
| Informational | No human interruption | Why the event enters the on-call path at all |
Choose actual time thresholds from business impact and escalation design. Do not copy a five-minute target from another company whose service, team size and consequences are unrelated to yours.
A weekly review should end with a change
Keep the review to thirty minutes and use a fixed agenda:
- Five minutes: raw alerts, groups and pages compared with the previous window.
- Five minutes: top noisy services and alert names.
- Five minutes: MTTA and resolution p50/p95, including data coverage.
- Five minutes: critical-open backlog, oldest items and current impact.
- Five minutes: escalations, after-hours load and handoff failures.
- Five minutes: choose one improvement with an owner and due date.
The output should look like this:
Observation: Checkout raw alerts rose 3.4x; groups stayed flat.
Evidence: DiskQueueDepth produced 68% of raw volume.
Action: Increase persistence window and fix duplicate dimensions.
Owner: Payments observability
Due: 2026-09-18
Success: raw volume down 60%, group detection unchanged.
“Continue monitoring” is not an action unless it names what will be monitored, by whom, until when and what result changes the decision. Otherwise it is the operational equivalent of putting the alert under a small blanket.
Use IncidentRelay Service Analytics as the starting point
Open Services → Analytics and choose a 7, 30, 90 or 365-day window. The view includes grouped-alert trends, raw alert volume, firing groups and a per-service table with open groups, critical-open groups, deduplication ratio, average MTTA/MTTR, maintenance activity and current impact context.
The API supports deeper sorting and automation:
GET /api/services/analytics
?days=30
&include_series=true
&include_noise=true
&include_response=true
&include_maintenance=true
&include_impact=true
&sort=raw_alerts
&order=desc
Useful sort keys include open_alert_groups, critical_open_alert_groups, raw_alerts, dedup_ratio, mtta, mttr and blast_radius. Start with the noisiest services, then switch to response tails and critical backlog. One giant dashboard is not required; a sequence of good questions works better.
A 30-day rollout
- Week 1: agree on units, metric definitions and two or three service classes. Record current data coverage.
- Week 2: baseline raw volume, groups, deduplication, MTTA/MTTR distributions and open backlog. Do not set targets yet.
- Week 3: investigate the top two noise sources and the slowest response tail. Ship one alerting or routing improvement.
- Week 4: add page and fairness data, set provisional targets and verify that no metric rewards hiding work.
After thirty days, keep only metrics that lead to a decision. If nobody can name what a widget changes, archive it. Dashboards also deserve retention policies.
The anti-gaming checklist
- Raw alerts, alert groups and human pages are reported separately.
- p50 and p95 are shown together with sample size and timestamp coverage.
- ACK means explicit human ownership, not automated metric cosmetics.
- Resolution is paired with recurrence and customer-impact context.
- Deduplication is reviewed for hidden noise and over-grouping.
- Current impact is not presented as historical impact.
- Maintenance suppression is reviewed beside total volume.
- Targets differ by severity and service class.
- Metrics evaluate systems and processes, not individual hero rankings.
- Every review produces one owned, measurable improvement.
Measure the system you want to improve
A good on-call scorecard tells a story: how much demand arrived, how effectively it was grouped, when a human accepted ownership, how long the alert remained unresolved and whether the burden was sustainable. No single number can carry that story alone.
Use metrics to find noisy sources, weak routing, slow tails and unfair load. Then change the system and watch whether the whole scorecard improves. If a number gets better while responders sleep less or customers wait longer, the number is not winning. It is merely escaping supervision.
For IncidentRelay's service model and analytics semantics, read Services, Impact and Analytics. The Services API guide documents filters, sort keys and the versioned analytics response.
Want to make this less theoretical?
IncidentRelay gives you schedules, rotations, routing, escalations and alert actions in one open-source, self-hosted package. The pager may still be rude, but at least it will be organized.