API Reference · A-0739–A-0740

Job Policies

How Priostack handles job deadlines, retries, and timer boundary events, and what to do in your model and worker to make service tasks resilient.

Overview· Workers· Incidents· Correlation IDs· Job Policies

Job deadlines

Job deadline field

Every activated job carries a deadline: an integer (JSON number) in Unix epoch milliseconds, never a date string. It is fixed by the server at activation:

deadline = activation time (Unix ms) + 300000   // 5 minutes

The deadline is informational only. There is no lock expiry: nothing happens when it passes. An activated job stays active until a worker completes or fails it, its instance is cancelled, or the server restarts (every outstanding job is then restored to waiting with the same key and retries). No second worker receives it in the meantime, and no incident is raised; the instance simply stays ACTIVE. GET /api/v1/jobs shows deadline as 0.

Field (activate request)TypeDefaultDescription
timeoutinteger (ms)noneAccepted and ignored. It sets neither a lock nor the deadline.

A job abandoned by a crashed worker is released only by POST /api/v1/jobs/{key}/fail with retries above 0 (it is re-queued at once) or by a server restart. To bound work in time, put the timeout in your worker, and model a timer boundary event on the task (below) when the process itself must move on:

// Bound your own processing time: the server never reclaims the job.
const timeout = new Promise((_, reject) =>
  setTimeout(() => reject(new Error("processing took too long")), 60_000));
try {
  const result = await Promise.race([processPayment(job), timeout]);
  await completeJob(job.key, result);
} catch (err) {
  // retries is stored as sent: one less uses up one attempt.
  await failJob(job.key, job.retries - 1, err.message);
}

Retry policy

Job retry configuration

Each job starts with the retries of its task: the retries attribute of zeebe:taskDefinition (an integer of 1 or more; retries="0" or a negative value is rejected at deploy with 422), or 3 when the attribute is absent. Every activated job carries its current retries.

When a worker calls fail, the body is {"retries": int, "errorMessage": string}. retries is the absolute number of attempts left: the server stores it as sent and does not decrement it. Send job.retries - 1 to use up one attempt, or job.retries unchanged to hand the job back without using one. An omitted retries counts as 0.

When retries reach 0 (or less), the job is removed, the instance state becomes INCIDENT, and an incident record (error_code JOB_FAILED, severity critical) appears in GET /api/v1/incidents. There is currently no REST call that retries or resumes an INCIDENT instance: PATCH /api/v1/incidents/{id} marks the record resolved and leaves the instance in INCIDENT.

ScenarioWorker actionResult
Transient error, more retries leftfail with retries: job.retries - 1 (above 0)Job re-queued immediately, with no delay
Hand the job back (shutdown)fail with retries: job.retriesJob re-queued immediately, no attempt used
Permanent errorfail with retries: 0Instance becomes INCIDENT; incident record created
Successcomplete (answers 204)Process advances to next element

Setting retries in BPMN (extension element)

A fragment: paste it inside a document that declares the bpmn namespace and xmlns:zeebe="http://camunda.org/schema/zeebe/1.0".

<bpmn:serviceTask id="task_charge" name="Charge payment">
  <bpmn:extensionElements>
    <zeebe:taskDefinition type="payment:charge" retries="3" />
  </bpmn:extensionElements>
</bpmn:serviceTask>

Retry delay

There is no retry delay on the server: a job failed with retries above 0 can be activated again immediately, and the fail body has no backoff field. When an external API needs time to recover, wait in the worker before handing the job back (the job stays yours while you wait), or model the pause in the process with a timer:

// The server has no retry delay, so the worker waits before failing.
const attempt = 3 - job.retries;            // with the default of 3 retries
await sleep(Math.pow(2, attempt) * 1000);   // 1 s, 2 s, 4 s
await failJob(job.key, job.retries - 1, "Stripe rate limit");

Timeout boundary events

BPMN timeout boundary events

A non-interrupting or interrupting timer boundary event on a service task lets you react to long-running jobs in the process model itself, without relying solely on job-level retries.

Interrupting vs non-interrupting: An interrupting boundary event cancels the task when it fires. A non-interrupting event triggers a parallel branch while the task continues running.

Common patterns:

PatternConfigurationUse case
Escalation after SLANon-interrupting timer → notification taskNotify a manager if payment hasn't completed in 30 min
Hard deadline cancelInterrupting timer → cancellation task → terminate end event (an error end event is read as a plain end here: it throws nothing)Cancel order if fulfilment doesn't start within 2 hours
Fallback pathInterrupting timer → alternative service taskSwitch to backup payment provider after 10 seconds

BPMN snippet - interrupting timer

A fragment, like the one above: paste it inside a document that declares the bpmn namespace.

<bpmn:boundaryEvent id="timeout_charge" attachedToRef="task_charge"
    cancelActivity="true">
  <bpmn:timerEventDefinition>
    <bpmn:timeDuration>PT30S</bpmn:timeDuration>
  </bpmn:timerEventDefinition>
</bpmn:boundaryEvent>

When an interrupting timer fires, the task's job is withdrawn: a later complete or fail on its key answers 404 {"error":"job not found"}. No incident is raised. The instance follows the sequence flow from the boundary event to your error-handling path. A non-interrupting timer leaves the job in place.

See also:   Worker API · Incident management · Troubleshooting guide