Job Policies
How Priostack handles job deadlines, retries, and timer boundary events, and what to do in your model and worker to make service tasks resilient.
Job deadlines
Job deadline field
Every activated job carries a deadline: an integer (JSON number) in Unix epoch milliseconds, never a date string. It is fixed by the server at activation:
deadline = activation time (Unix ms) + 300000 // 5 minutes
The deadline is informational only. There is no lock expiry: nothing happens when it passes. An activated job stays active until a worker completes or fails it, its instance is cancelled, or the server restarts (every outstanding job is then restored to waiting with the same key and retries). No second worker receives it in the meantime, and no incident is raised; the instance simply stays ACTIVE. GET /api/v1/jobs shows deadline as 0.
| Field (activate request) | Type | Default | Description |
|---|---|---|---|
timeout | integer (ms) | none | Accepted and ignored. It sets neither a lock nor the deadline. |
A job abandoned by a crashed worker is released only by POST /api/v1/jobs/{key}/fail with retries above 0 (it is re-queued at once) or by a server restart. To bound work in time, put the timeout in your worker, and model a timer boundary event on the task (below) when the process itself must move on:
// Bound your own processing time: the server never reclaims the job.
const timeout = new Promise((_, reject) =>
setTimeout(() => reject(new Error("processing took too long")), 60_000));
try {
const result = await Promise.race([processPayment(job), timeout]);
await completeJob(job.key, result);
} catch (err) {
// retries is stored as sent: one less uses up one attempt.
await failJob(job.key, job.retries - 1, err.message);
}
Retry policy
Job retry configuration
Each job starts with the retries of its task: the retries attribute of zeebe:taskDefinition (an integer of 1 or more; retries="0" or a negative value is rejected at deploy with 422), or 3 when the attribute is absent. Every activated job carries its current retries.
When a worker calls fail, the body is {"retries": int, "errorMessage": string}. retries is the absolute number of attempts left: the server stores it as sent and does not decrement it. Send job.retries - 1 to use up one attempt, or job.retries unchanged to hand the job back without using one. An omitted retries counts as 0.
When retries reach 0 (or less), the job is removed, the instance state becomes INCIDENT, and an incident record (error_code JOB_FAILED, severity critical) appears in GET /api/v1/incidents. There is currently no REST call that retries or resumes an INCIDENT instance: PATCH /api/v1/incidents/{id} marks the record resolved and leaves the instance in INCIDENT.
| Scenario | Worker action | Result |
|---|---|---|
| Transient error, more retries left | fail with retries: job.retries - 1 (above 0) | Job re-queued immediately, with no delay |
| Hand the job back (shutdown) | fail with retries: job.retries | Job re-queued immediately, no attempt used |
| Permanent error | fail with retries: 0 | Instance becomes INCIDENT; incident record created |
| Success | complete (answers 204) | Process advances to next element |
Setting retries in BPMN (extension element)
A fragment: paste it inside a document that declares the bpmn namespace and xmlns:zeebe="http://camunda.org/schema/zeebe/1.0".
<bpmn:serviceTask id="task_charge" name="Charge payment">
<bpmn:extensionElements>
<zeebe:taskDefinition type="payment:charge" retries="3" />
</bpmn:extensionElements>
</bpmn:serviceTask>
Retry delay
There is no retry delay on the server: a job failed with retries above 0 can be activated again immediately, and the fail body has no backoff field. When an external API needs time to recover, wait in the worker before handing the job back (the job stays yours while you wait), or model the pause in the process with a timer:
// The server has no retry delay, so the worker waits before failing. const attempt = 3 - job.retries; // with the default of 3 retries await sleep(Math.pow(2, attempt) * 1000); // 1 s, 2 s, 4 s await failJob(job.key, job.retries - 1, "Stripe rate limit");
Timeout boundary events
BPMN timeout boundary events
A non-interrupting or interrupting timer boundary event on a service task lets you react to long-running jobs in the process model itself, without relying solely on job-level retries.
Common patterns:
| Pattern | Configuration | Use case |
|---|---|---|
| Escalation after SLA | Non-interrupting timer → notification task | Notify a manager if payment hasn't completed in 30 min |
| Hard deadline cancel | Interrupting timer → cancellation task → terminate end event (an error end event is read as a plain end here: it throws nothing) | Cancel order if fulfilment doesn't start within 2 hours |
| Fallback path | Interrupting timer → alternative service task | Switch to backup payment provider after 10 seconds |
BPMN snippet - interrupting timer
A fragment, like the one above: paste it inside a document that declares the bpmn namespace.
<bpmn:boundaryEvent id="timeout_charge" attachedToRef="task_charge"
cancelActivity="true">
<bpmn:timerEventDefinition>
<bpmn:timeDuration>PT30S</bpmn:timeDuration>
</bpmn:timerEventDefinition>
</bpmn:boundaryEvent>
When an interrupting timer fires, the task's job is withdrawn: a later complete or fail on its key answers 404 {"error":"job not found"}. No incident is raised. The instance follows the sequence flow from the boundary event to your error-handling path. A non-interrupting timer leaves the job in place.