Incident Management API
An incident record is created automatically when a worker fails a job with no retries left. Use these endpoints to list, create, resolve and export incident records.
INCIDENT is a process instance state: the engine sets it whenever a run fails and cannot move on (a job failed with no retries left, a decision that cannot be evaluated, a gateway condition that cannot be read, a timer error). An incident record, listed below, is created only for the first case, a job failed with retries of 0 or less. To find every failed run, including those with no record, use GET /api/v1/process-instances?filter.state=INCIDENT.
Returns every incident record of your account, open and resolved, in one response. There is no pagination and there are no filters: query parameters are ignored, and the order is not specified. An X-Correlation-ID request header is echoed back.
Incident fields
| Field | Type | Description |
|---|---|---|
id | string | inc- followed by 32 hex characters. |
instance_key | string | The process instance key, e.g. run_1ce49bbc5260b39423bd. |
process_key | string | The instance's processDefinitionKey. |
element_id | string | For a record the engine created, the key of the failed job. |
error_code | string | JOB_FAILED for a record the engine created. |
error_message | string | The errorMessage the worker sent with its last fail call. |
severity | string | critical, warning or info. |
created_at, resolved_at | string | RFC 3339 times; resolved_at only once resolved. |
api_key | string | Your own key, echoed in full. Do not forward these records to third parties as they are. |
Optional fields appear when set: resolved_by, root_cause, post_mortem_url, tags, retry_count, alert_sent, correlation_id.
Example response 200 OK
{
"items": [
{
"id": "inc-4f1c0a9b2d7e4e0c8a51b6f3d2c9e7a1",
"api_key": "ps_your_key",
"process_key": "3f9c2a7e5b1d8c4e6a02",
"instance_key": "run_1ce49bbc5260b39423bd",
"element_id": "a99f579d62ae71f8764f",
"error_code": "JOB_FAILED",
"error_message": "Stripe timeout after 3 retries",
"severity": "critical",
"created_at": "2026-09-29T09:05:00Z"
}
],
"total": 1,
"open_count": 1
}
Records an incident of your own, for example one your worker detected. Answers 201 Created with the record.
Request body
| Field | Type | Description |
|---|---|---|
instance_key, process_key, element_id | string | What the incident is about. |
error_code, error_message | string | Your own code and message. |
severity | string | critical, warning (default) or info. |
tags | array of strings | Optional labels. |
retry_count | integer | Optional. |
Marks the record resolved. It does not touch the process instance, which stays in INCIDENT. A JSON body is required; {} is enough.
Request body
| Field | Type | Description |
|---|---|---|
resolved_by | string | Who resolved it. Defaults to your key. |
root_cause | string | Optional note. |
post_mortem_url | string | Optional link. |
Example response 200 OK
{ "id": "inc-4f1c0a9b2d7e4e0c8a51b6f3d2c9e7a1", "status": "resolved" }
An unknown id (or another account's) answers 404; a record already resolved answers 422.
Body {"ids": ["inc-...", "inc-..."], "resolved_by": "..."}. Answers {"resolved": 2, "status": "ok"}.
Returns every record of your account as text/csv, without the api_key column.
Retries are decided when the worker fails the job: POST /api/v1/jobs/{key}/fail with retries above 0 puts the job back in the queue at once, and 0 or less raises the incident. Once an instance is in INCIDENT, no REST route resumes or cancels it today, and there is no route to change a job's retries. Resolving the record closes the record only.
No webhook event is fired for incidents (see the event list). To alert on them, poll GET /api/v1/incidents and watch open_count, or poll GET /api/v1/process-instances?filter.state=INCIDENT.