This is a working outline, not the finished essay. Sections and the key points to hit are laid out;
the prose goes on top. Note on the title number: the suite currently runs 124 scheduled
jobs total (verified via
crontab -l); pin the exact trading-scanner subset before publishing if you want a number in the title.
1. The hook: "it ran" is not "it worked"
Open on the core insight that reframes the whole piece.
- A cron job exiting 0 tells you the process ran — not that the data it produced is correct.
- In a trading pipeline the scary failures are silent: the job succeeds, the data is subtly wrong, and the error only surfaces downstream as a corrupted statistic.
- Thesis: on self-hosted infrastructure, your real job is detecting correctness failures, not crashes.
2. The shape of the pipeline
Give the reader the map before the war stories.
- The daily arc: overnight price load → close-of-day scanners → scoring → static export + deploy → alerts.
- Why it's many small scheduled jobs instead of one monolith: isolation, independent retries, clear logs per stage.
- The dependency problem: a 4 PM scan is worthless if the 3 AM load silently half-finished.
3. Failure mode: silent data gaps
The flagship war story — lead with the most instructive one.
- The incident: a multi-week price gap that went unnoticed and quietly corrupted accuracy statistics until a calibration looked wrong.
- Why it was invisible: every job "succeeded"; the data just stopped arriving for a slice of the universe.
- The lesson that changed the design: check date continuity before you trust any statistic computed from the data.
4. Failure mode: partial loads and stale tickers
The everyday corruption, less dramatic but more frequent.
- A load that pulls most of the universe but drops a slice — totals look plausible, so nothing screams.
- Dead/delisted tickers that stop updating and pollute signals if not flagged.
- The fix: a stale-price report that flags any symbol whose data stops advancing.
5. Failure mode: the calendar and the clock
The boring bugs that cost the most time.
- Market holidays and half-days: a "missing" day that's actually correct, and how to distinguish it from a real gap.
- Timezone/DST drift: a job scheduled in local time that silently shifts relative to the market open twice a year.
- Overlapping runs and long jobs bleeding into the next window.
6. Failure mode: the crontab itself
A genuinely non-obvious infra trap worth its own section.
- cron's minimal environment: the path/temp-dir assumptions that hold in your shell and break under cron.
- A real one: a long temp-directory path truncating and breaking job installation — fixed by controlling
TMPDIRat install time. - Why the crontab is itself critical state: back it up, version it, verify job count after every change.
7. The monitoring that actually catches this
The payoff — concrete, cheap, self-hosted observability.
- A date-continuity monitor that alerts on missing trading days or partial loads across the whole universe.
- A scanner-health view: last-run timestamps and candidate counts per scanner, so a dead scanner is visible at a glance.
- Push alerts (Telegram) for the few conditions that need a human now — and the discipline of alerting on correctness, not just failure.
- A one-command health check: DB freshness, data-file ages, service status, job count — the first thing you run when something smells off.
8. Principles that generalize
Lift the specifics into reusable advice.
- Monitor outputs, not just exit codes.
- Make "no data" an alarm, not a silent zero.
- Treat your scheduler config as production state.
- Backups and reproducibility beat cleverness when it's 3 AM and a load failed.
Takeaway to land: running it yourself teaches you that reliability is mostly about noticing the quiet wrongness fast.