Process Automation · Reliability · Production · Scheduling

Twice in the same hour, or not at all: what the crontab does to a daily pipeline

Published · Updated

A pipeline that publishes every day without supervision rests on one assumption: the trigger fires exactly once. The cron manual page is explicit that during clock conversions, it does not.

A daily automation has one load-bearing assumption underneath it: the schedule fires once per day. The Tatano Energy platform publishes 8 SEO articles every day with 0 manual intervention, across 4 separately indexed country domains and 7 languages. Nobody is watching the trigger. That makes it worth reading what the scheduler actually promises, because the crontab manual page is unusually candid about the conversion windows where it promises nothing at all.

The clock is not a guarantee

The crontab(5) page states that non-existent times, such as the missing hours during the daylight savings time conversion, will never match, causing jobs scheduled during the missing times not to be run. It also states that times that occur more than once during the conversion will cause matching jobs to be run twice. So a job pinned inside the shifted window either vanishes for a day or executes twice, and which one you get depends on the direction of the shift.

The same page documents CRON_TZ, which specifies the time zone specific for the cron table. That is the lever worth pulling first. A schedule interpreted in a zone that does not observe the shift has no missing hour and no repeated hour, so the ambiguity disappears at the source rather than being absorbed downstream. If local wall-clock time genuinely matters for the business, move the job out of the shifted window instead.

A missing newline decides whether anything runs

The quieter failure sits in the same document: if the last entry in a crontab is missing a newline, terminated by EOF, cron will consider the crontab at least partially broken. No retry masks that, and no error surfaces inside your application, because the application was never started. An automation whose selling point is 0 manual intervention is also an automation where nobody notices silence, so the trigger layer needs a check of its own.

Duplicate trigger, single claim

Hardening the schedule reduces duplicate runs but cannot make them impossible, so the real boundary belongs one layer down, where work is claimed. PostgreSQL documents that with SKIP LOCKED, any selected rows that cannot be immediately locked are skipped, and notes this is not suitable for general-purpose work because it provides an inconsistent view of the data, but can be used to avoid lock contention with multiple consumers accessing a queue-like table.

That inconsistent view is exactly what you want when the same job starts twice. The second starter finds the row already locked and walks past it rather than queuing behind it. NOWAIT is the other posture documented there: the statement reports an error, rather than waiting, if a selected row cannot be locked immediately. Choose skipping when a duplicate should quietly do nothing, and an error when a duplicate should be loud.

A double run is a positive feedback candidate

Duplicate work is not only wasted work. The SRE book defines a cascading failure as a failure that grows over time as a result of positive feedback, and notes that if retries are not randomly distributed over the retry window, a small perturbation such as a network blip can cause retry ripples to schedule at the same time. Two synchronised runs hitting the same downstream service are precisely that kind of perturbation.

Two mitigations from the same chapter apply directly. Always use randomised exponential backoff when scheduling retries, so duplicated work spreads instead of stacking. And consider a server-wide retry budget: the worked example allows only 60 retries per minute in a process, and when the budget is exceeded, the request simply fails rather than being retried. A budget turns an unbounded amplification into a bounded one, which is the difference between a bad hour and an outage.

Watch the trigger, not just the service

The four golden signals of monitoring are latency, traffic, errors and saturation. On a scheduled pipeline, traffic is the signal that catches a schedule that never fired: the absence of a run is invisible to every error-based alert. White-box monitoring therefore allows detection of imminent problems, failures masked by retries, and so forth, which is the category a job that silently ran twice and half-failed belongs to.

Resist over-probing in compensation. The SRE book notes that for a web service targeting no more than 9 hours aggregate downtime per year (99.9% annual uptime), probing for a 200 status more than once or twice a minute is probably unnecessarily frequent. The matching rule for alerts: if a page merely merits a robotic response, it shouldn't be a page. A duplicate run that SKIP LOCKED already absorbed does not need to wake anyone.

The practical sequence is cheap to adopt. Pin the schedule to a zone without a conversion window, or move the job away from the shifted hour. Validate that the crontab file ends with a newline, because the whole table depends on it. Claim work through a locking clause so a second start is a no-op rather than a second run. Then alert on missing traffic, not only on errors, because the failure mode of a silent pipeline is silence.

Sources

man7.org (cronie) — crontab(5) Linux manual page — https://man7.org/linux/man-pages/man5/crontab.5.html

PostgreSQL 18 Documentation — SELECT (The Locking Clause) — https://www.postgresql.org/docs/current/sql-select.html

Google SRE Book — Addressing Cascading Failures — https://sre.google/sre-book/addressing-cascading-failures/

Google SRE Book — Monitoring Distributed Systems — https://sre.google/sre-book/monitoring-distributed-systems/

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Working on a project where these methods apply?