Automation · Custom Software

SKIP LOCKED or NOWAIT: how parallel workers claim rows from one queue table

Published · Updated

Fan-out is easy to draw and awkward to implement. The hard part is the claim: how one worker takes a row of work without a second worker taking the same one. Postgres documents two answers.

Fan-out in a daily pipeline is easy to draw: one scheduled trigger, a set of work items, several workers. The awkward part is never the parallelism. It is the claim, meaning how one consumer takes a row without a second consumer taking the same row. PostgreSQL documents two different answers to that question, and they disagree about what should happen when the row you selected is already locked by someone else.

The workload decides how much the claim matters

For Tatano Energy we built a multilingual e-commerce platform deployed across four country domains, with a daily SEO autoblog driven by search trends. The published results: country domains indexed separately, 4; languages served, 7; SEO articles published every day, 8; manual intervention required, 0. That last figure is the one that constrains the design. Nobody is watching to notice that two workers produced the same article, so the claim has to be unambiguous without a human in the loop.

SKIP LOCKED buys throughput with an inconsistent view

The PostgreSQL documentation is direct about the trade. With SKIP LOCKED, any selected rows that cannot be immediately locked are skipped. It also states that skipping locked rows provides an inconsistent view of the data, so this is not suitable for general-purpose work, but can be used to avoid lock contention with multiple consumers accessing a queue-like table. Read that as permission with a boundary: it is the right tool for claiming queue rows and the wrong tool for reporting on them.

NOWAIT answers a different question

NOWAIT behaves differently: the statement reports an error, rather than waiting, if a selected row cannot be locked immediately. That is what you want when a specific item must be handled by this particular caller, and contention is a fault you would rather surface than route around. SKIP LOCKED says give me any available work. NOWAIT says give me this work or tell me now. Picking the wrong one turns a scheduling detail into either silent duplication or needless errors.

Fan-in cannot be read off the queue

The consequence of skipping is that a single query over the queue is no longer an inventory. Rows that are locked and in progress look, to that query, exactly like rows that are not there. So the fan-in side, deciding that the day's batch is complete, has to be derived from durable per-item state rather than from one observation of the table. If a pipeline runs with zero manual intervention, that completion decision belongs to the system, not to whoever glances at it.

The claim starts at the scheduler

A fan-out that begins with cron inherits cron's calendar quirks. The crontab manual notes that non-existent times, such as the missing hours during the daylight savings time conversion, will never match, causing jobs scheduled during the missing times not to be run, and that times that occur more than once during the same conversion will cause matching jobs to be run twice. It also documents CRON_TZ for the table's time zone, and warns that a crontab whose last entry lacks a newline is considered at least partially broken.

Retries multiply fan-out

Every skipped or failed item becomes a retry, and retries are themselves a fan-out. The SRE guidance is that a cascading failure is a failure that grows over time as a result of positive feedback, and that you should always use randomised exponential backoff when scheduling retries. Without randomisation, a small perturbation such as a network blip can cause retry ripples to schedule at the same time. The same source suggests a server-wide retry budget: only allow 60 retries per minute in a process, and fail the request once that is exceeded.

Watch saturation, not just success

Four signals cover the queue: latency, traffic, errors and saturation. Saturation is the one that tells you workers are contending rather than working. White-box monitoring matters here because it allows detection of imminent problems and failures masked by retries, which is precisely how a duplicated or starved queue hides. On alerting, the same source is blunt: if a page merely merits a robotic response, it shouldn't be a page, and for a service targeting 99.9% annual uptime, probing for a 200 status more than once or twice a minute is probably unnecessarily frequent.

If you are building a fan-out pipeline, write down three decisions before the code: which locking clause claims a row and why, where completion state lives so fan-in does not depend on reading the queue, and what the retry budget is. Then check the scheduler's time zone and the newline at the end of the crontab. Those are small, cheap choices that determine whether unattended runs stay unattended.

Sources

PostgreSQL 18 Documentation — SELECT (The Locking Clause) — https://www.postgresql.org/docs/current/sql-select.html

man7.org (cronie) — crontab(5) Linux manual page — https://man7.org/linux/man-pages/man5/crontab.5.html

Google SRE Book — Addressing Cascading Failures — https://sre.google/sre-book/addressing-cascading-failures/

Google SRE Book — Monitoring Distributed Systems — https://sre.google/sre-book/monitoring-distributed-systems/

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Working on a project where these methods apply?