Incidents
Incidents group related topology events (e.g., a flapping link producing multiple link_state_changed events) into a single correlated entity. All authenticated users can list, view, acknowledge, and resolve incidents.
List Incidents
GET /api/v1/incidents
GET /api/v1/incidents?area_id={areaID}&status={status}&from={from}&to={to}&limit=100&offset=0
List incidents with optional filters. Returns newest first.
Query parameters (all optional):
area_id -- Filter by area UUID
status -- Filter by incident status ("active", "acknowledged", "resolved")
from -- Start time filter, RFC3339 (e.g., "2026-02-16T00:00:00Z")
to -- End time filter, RFC3339 (e.g., "2026-02-17T00:00:00Z")
limit -- Maximum number of incidents to return (default: 100)
offset -- Number of incidents to skip for pagination (default: 0)
Response: 200 OK
{
"incidents": [
{
"id": "uuid",
"area_id": "uuid",
"status": "active",
"severity": "warning",
"root_cause_type": "device_down",
"root_cause_id": "10.0.0.1",
"summary": "Router 10.0.0.1 failure: 3 links down, 2 adjacencies lost",
"event_count": 5,
"first_event_at": "2026-02-17T14:30:00Z",
"last_event_at": "2026-02-17T14:30:02Z",
"detail": {
"root_cause_type": "device_down",
"affected_routers": ["10.0.0.1", "10.0.0.2"]
},
"created_at": "2026-02-17T14:30:05Z"
}
],
"total": 12,
"limit": 100,
"offset": 0
}
Notes:
- The
total field contains the total count of incidents matching the filter (before limit/offset), for pagination.
resolved_at, acknowledged_at, acknowledged_by, and root_cause_id are omitted when unset.
root_cause_type values: device_down, node_failure, asbr_withdrawal, unknown (event-correlated), and area_partition (detector-driven, see below).
area_partition incidents (RFC 2328 §3.7) are raised by the engine's partition detector, not by event correlation: event_count is 0, root_cause_id is absent, and detail carries the split — area_label, consequence ("bridged" | "isolated" | "backbone_partition" | "unknown"), and components, an array of islands each with collector_ids, collector_names, router_ids, and has_backbone_abr. They auto-resolve on positive evidence; detail.resolution_reason records why (e.g. "recorder views overlap again"). Severity follows the consequence: bridged/unknown → warning, isolated/backbone_partition → critical.
Error responses:
400 Bad Request -- Invalid UUID, invalid RFC3339 timestamp, or invalid status value
401 Unauthorized -- Missing or invalid access token
Get Incident
GET /api/v1/incidents/{incidentID}
Get a single incident by ID, including its child events and display names.
Response: 200 OK
{
"id": "uuid",
"area_id": "uuid",
"status": "active",
"severity": "warning",
"root_cause_type": "device_down",
"root_cause_id": "10.0.0.2",
"summary": "Router 10.0.0.2 failure: 3 links down",
"event_count": 3,
"first_event_at": "2026-02-17T14:30:00Z",
"last_event_at": "2026-02-17T14:35:00Z",
"detail": {
"root_cause_type": "device_down",
"affected_routers": ["10.0.0.2", "10.0.0.3"]
},
"created_at": "2026-02-17T14:30:05Z",
"events": [
{
"id": "uuid",
"event_time": "2026-02-17T14:35:00Z",
"area_id": "uuid",
"collector_id": "collector-nyc-dc1-01",
"event_type": "link_state_changed",
"entity_type": "link",
"entity_id": "uuid",
"router_id": "10.0.0.2",
"detail": {
"router_a": "10.0.0.2",
"router_b": "10.0.0.3",
"old_state": "up",
"new_state": "down"
},
"incident_id": "uuid"
}
],
"display_names": {
"10.0.0.2": "core-rtr-01.ams",
"10.0.0.3": "edge-rtr-02.ams"
}
}
Notes:
- The
events array contains the child topology events belonging to this incident, ordered by event_time descending (newest first).
- The
display_names field maps router IDs found across all child events to human-readable device names (hostname preferred, dns_name as fallback). Omitted when no names are available.
Error responses:
404 Not Found -- Incident does not exist
401 Unauthorized -- Missing or invalid access token
Acknowledge Incident
PUT /api/v1/incidents/{incidentID}/acknowledge
Acknowledge an open incident. Sets acknowledged_at to the current time and records the acknowledging user's ID in acknowledged_by. No request body required.
Response: 200 OK (updated incident object with status: "acknowledged", acknowledged_at and acknowledged_by set)
Error responses:
401 Unauthorized -- Missing or invalid access token
404 Not Found -- Incident does not exist
500 Internal Server Error -- Database error
Resolve Incident
PUT /api/v1/incidents/{incidentID}/resolve
Manually resolve an incident. Sets status to "resolved" and resolved_at to the current time. No request body required. (area_partition incidents also resolve automatically when the partition heals — see List Incidents notes.)
Response: 200 OK (updated incident object with status: "resolved" and resolved_at set)
Error responses:
401 Unauthorized -- Missing or invalid access token
404 Not Found -- Incident does not exist
500 Internal Server Error -- Database error