Skip to content
// jays.website / trading / case study

Trading Intelligence — a scanner that runs itself

A self-hosted quantitative research platform that pulls the whole US equity/ETF market into a database every night, runs twelve independent scanners over it, gates the results against the market regime, and scores what survives — then ships everything as static files. No server to babysit; the output is a git commit.

~2,350 symbols nightly 12 versioned scanners 124 scheduled jobs ~1.2M daily bars stored
The problem

Consistency is impossible to do by hand

Scanning thousands of symbols every single day for a handful of high-conviction setups is exactly the kind of work a human does well for a week and then stops doing. Attention drifts, the universe changes, and — worst of all — there's no honest record of whether the picks were any good. I wanted a system that would do the scan the same way every day, keep score on itself, and suppress its own output when the market made its signals untrustworthy.

The constraints were self-imposed and deliberate: run it on hardware I already own, spend nothing on hosting, and make every day's opinion auditable after the fact.

Approach

One database, many scanners, one referee

Rather than one clever model, the design is a pipeline of simple, measurable stages. A single SQLite database is the one source of truth. Each scanner is an independent, versioned strategy that reads from it — so a scanner's measured accuracy belongs to a specific version, not a blurred average across rewrites. A market-regime model acts as a referee that can veto output wholesale. Conviction scoring then weighs what's left. Everything downstream is a pre-built static file, which is what makes the whole thing cheap and crash-proof.

Architecture

From nightly bars to a conviction-scored pick

Four tiers run end to end, unattended, every market day — each stage a scheduled job with its own logging:

Trading intelligence pipeline architecture Four tiers. Ingest: Polygon.io nightly bars, SEC EDGAR filings for DCF, and 13F institutional holdings feed a central SQLite database. Compute: twelve versioned scanners read the database, their output passes through a market-regime gate, then a conviction-scoring layer. Deliver: results become static HTML and JSON on GitHub Pages, Telegram alerts, public paper portfolios, and a Claude-written top-25 plus weekly newsletter. INGEST STORE COMPUTE DELIVER Polygon.io nightly grouped bars SEC EDGAR filings → DCF model 13F holdings institutional flow SQLite central DB — one source of truth 12 versioned scanners regime gate veto in bad tapes conviction scoring Static HTML + JSON → GitHub Pages Telegram alerts high-conviction only Paper portfolios marked to market Claude top-25 + weekly newsletter
Nightly pipeline: ingest → store → compute → deliver. Every arrow is a scheduled, logged job.

Want the subsystem-by-subsystem tour — the data hygiene, the scanner families, how breadth feeds the regime model? That lives on the trading architecture deep-dive.

Key technical decisions

What I chose, and what it cost

Static-first delivery

No application server sits in the read path. Every output is a pre-built HTML/JSON file committed to git and served from the edge.

Tradeoff The site can't crash and costs ~nothing, but data is end-of-day — there's no live intraday interactivity in the public view.

A regime gate that says "no"

A daily breadth-and-trend model classifies the tape. In hostile regimes the pipeline suppresses signal output entirely.

Tradeoff Fewer signals, and some real bear-market bounces get skipped — accepted, because the alternative is confidently shipping fair-weather noise.

Versioned scanners, measured honestly

Accuracy is tracked per scanner version. Rewrites start a new track record instead of inheriting an old one.

Tradeoff It takes longer to trust a new scanner, but the numbers never flatter a strategy that was quietly rewritten.

Confluence over cleverness

A signal confirmed by several independent scanners outranks any single-scanner hit. The edge is in the stack, not one indicator.

Tradeoff The highest-conviction list is short by design — breadth of coverage is traded for precision.

Validate over a full cycle

A hard-won lesson: edges measured only in a correction reversed in the following bull. Backtests now have to survive a full market cycle before they change a weight.

Tradeoff Promising short-window results sit on the bench longer before they're allowed to influence capital.

AI ranks; math decides

A Claude analyst writes the nightly ranked top-25 with reasoning, but the conviction score itself stays deterministic and reproducible.

Tradeoff The narrative layer can't be blindly trusted — it's explicitly kept out of the scoring that drives paper-trade entries.
Outcomes

What it actually does, every day

~2,350
symbols pulled nightly
src: daily_prices (DB)
12
active scanners
src: scanners.json
124
scheduled cron jobs
src: crontab
~1.2M
daily bars in store
src: DB (~900 MB)
  • Runs unattended. The full ingest → score → deliver cycle completes on schedule after each close with no manual step.
  • Keeps score in public. Top-conviction picks enter paper portfolios automatically, carry hard stop- and time-losses, and resolve to real closed outcomes — no post-hoc cherry-picking.
  • Alerts only when it matters. Telegram fires for high-conviction setups the moment scoring finishes; everything else stays on the dashboard.
  • Versioned opinions. Every day's output is a git commit, so the entire history of what the system believed is reconstructable.

Honest note on "signals tracked": earlier versions of this site cited 100k–264k tracked signals. The current longitudinal signal-to-outcome store is far thinner than that (hundreds of backtested records), so I've dropped the big number rather than repeat one I can't reproduce from the database today. Rebuilding that store properly is the first item below.

What I'd do next

The honest backlog

  • Rebuild the signal-tracking store. The conviction weights deserve a complete, longitudinal record of every emitted signal and its forward outcome — not the partial backtest table that exists now. This is the highest-value gap.
  • Move the authenticated dashboard off localhost. The public, static views are on the edge; the gated Flask app still runs locally pending migration to a free-tier Oracle A1.Flex instance.
  • Paper → small live fills. A live A/B of exit rules (tight vs. wide stops) is running; once it resolves, graduate the best-performing book to small real positions under the same discipline.
  • Broaden the universe. Today it's US equities and ETFs. A parallel crypto track already runs the same logic on a separate universe — the template for adding more.

See it live, or read the deeper engineering tour.