Static Documentation

RMON

Version latest · Updated 2026-06-22
Interactive docs View on GitHub

RMON integration#

IncidentRelay can receive alerts from RMON through a dedicated inbound integration. RMON sends each alert to IncidentRelay with a route intake token, and IncidentRelay processes it through the standard routing, grouping, suppression, notification, escalation, and explain-trace pipeline.

Endpoint#

POST /api/integrations/rmon

The endpoint requires the intake token of an active IncidentRelay route whose source is rmon.

Authorization: Bearer <route-intake-token>
Content-Type: application/json

Create an IncidentRelay route#

  • Open Routes in IncidentRelay.
  • Create a route or edit an existing one.
  • Select RMON as the source.
  • Select the team that owns the alerts.
  • Configure route matchers.
  • Configure grouping.
  • Select a rotation or escalation policy.
  • Make sure the route is active.
  • Copy the route intake token.

Recommended grouping:

[
  "rmon_check_id",
  "rmon_check_type"
]

Example matchers:

{
  "environment": "production",
  "rmon_region": "eu-west"
}

Only use stable labels for grouping. Do not group by values that change between repeated executions of the same check.

Configure RMON#

In the RMON IncidentRelay channel, set:

Channel name: https://incidentrelay.example.com
Token: <route-intake-token>

The channel URL must contain only the base IncidentRelay URL. RMON appends the integration path automatically:

/api/integrations/rmon

Do not place the token in the URL.

Alert lifecycle#

RMON sends:

firing

for active failures and:

resolved

for recovery events.

IncidentRelay updates the existing alert when the firing and resolved notifications use the same fingerprint.

Resolved-like values such as the following are also normalized to resolved:

clear
cleared
closed
info
normal
ok
recover
recovered
resolved
up

Other values are treated as firing.

Payload fields#

Required field#

FieldDescription
titleHuman-readable alert title

Alert identity#

FieldDescription
fingerprintStable key used for deduplication
external_idExternal check or event identifier
check_idRMON check identifier
multi_check_idRMON multi-check identifier
state_idRMON check-state identifier

Check context#

FieldDescription
rmon_nameName of the RMON instance
check_nameHuman-readable check name
check_typeCheck type, such as http
targetChecked host, URL, or other target
agentAgent that executed the check
regionCheck region
countryCheck country
descriptionCheck description

Routing and grouping#

FieldDescription
teamOptional IncidentRelay team slug fallback
severityAlert severity
statusAlert lifecycle status
labelsAdditional routing and grouping labels
FieldDescription
runbookRunbook text or URL
runbook_urlRunbook URL
event_linkDirect source event or check URL

Additional fields are accepted and retained in the stored payload.

Normalized labels#

IncidentRelay adds the following labels when corresponding values are available:

IncidentRelay labelSource field
alertnamecheck_name or title
severityseverity
rmon_namermon_name
rmon_check_idmulti_check_id, check_id, or existing label
rmon_state_idstate_id
rmon_check_namecheck_name
rmon_check_typecheck_type
targettarget or instance
rmon_agentagent
rmon_regionregion
rmon_countrycountry
runbook_urlrunbook_url
event_linkFirst available event or runbook link

Existing labels supplied by RMON are preserved.

Team selection#

IncidentRelay selects the optional team hint from the first available value:

  • top-level team;
  • labels.team;
  • labels.oncall_team.

The route remains authoritative. A team hint does not bypass route token, source, matcher, or access checks.

Severity#

IncidentRelay selects severity from:

  • top-level severity;
  • labels.severity;
  • warning.

Deduplication#

When fingerprint is present and valid, IncidentRelay uses it as the deduplication key.

Values such as the following are ignored:

None
null
None None
null null

When a valid fingerprint is unavailable, IncidentRelay generates a stable key from:

  • source rmon;
  • external or check identifier;
  • title;
  • normalized labels.

RMON should keep the same fingerprint between the firing and recovery notifications for one logical check failure.

Example payload#

{
  "title": "[rmon-production] critical: HTTP check failed",
  "message": "critical: HTTP check failed",
  "severity": "critical",
  "status": "firing",
  "fingerprint": "42 7",
  "external_id": 42,
  "rmon_name": "rmon-production",
  "multi_check_id": 42,
  "state_id": 7,
  "check_name": "Public API",
  "check_type": "http",
  "target": "https://api.example.com/health",
  "agent": "eu-west-agent",
  "region": "eu-west",
  "country": "DE",
  "runbook_url": "https://example.com/runbooks/public-api",
  "labels": {
    "team": "sre",
    "environment": "production"
  }
}

Manual test#

Save the example payload as rmon-payload.json and run:

curl -X POST   "https://incidentrelay.example.com/api/integrations/rmon"   -H "Authorization: Bearer ROUTE_INTAKE_TOKEN"   -H "Content-Type: application/json"   --data-binary @rmon-payload.json

A successful response contains one ingest result:

[
  {
    "created": true,
    "alert_id": 123,
    "group_id": 45,
    "status": "firing",
    "team_id": 2,
    "team_slug": "sre",
    "route_id": 7,
    "routing_error": null,
    "trace_id": "..."
  }
]

HTTP responses#

StatusMeaning
200Alert processed successfully
202Processing accepted but not fully completed synchronously
207Mixed processing outcome
400Invalid payload or routing failure
401Route intake token is missing or invalid

A routing failure response includes a trace_id. Administrators can inspect the explain trace to see route matcher evaluation and the exact failure reason.

Troubleshooting#

Route intake token is required#

RMON did not send the bearer token, or the channel token is empty.

Check the RMON channel configuration:

Token: <route-intake-token>

Alert did not match any active route#

Verify that:

  • the route is active;
  • the route source is rmon;
  • the token belongs to that route;
  • every configured matcher matches the received labels;
  • the team is active;
  • the route has a rotation or escalation policy.

Use the returned trace_id to inspect routing.

Recovery creates another alert#

Make sure the same fingerprint is sent for both firing and resolved notifications.

The fingerprint must identify the logical check failure and must not contain changing timestamps or random values.

Alerts from different checks are merged#

Use a fingerprint containing the check identifier and state identifier, or another stable combination unique to the logical check event.

Recommended example:

<multi_check_id> <state_id>

Set runbook_url to an absolute http:// or https:// URL. A non-URL runbook value remains available in the stored payload but is not used as a clickable event link.

Test notification behaves differently#

Test notifications may not contain state_id, check metadata, or production labels. IncidentRelay generates a fallback deduplication key when the test payload has no valid fingerprint.

Compare the labels shown in the explain trace with the route matchers.

Security#

  • Use HTTPS.
  • Keep the route intake token secret.
  • Use a separate route token for each RMON instance or trust boundary.
  • Rotate exposed tokens immediately.
  • Do not send the token inside the JSON payload.
  • Do not put the token in the URL.
  • Restrict route matchers to the expected RMON labels.