Skip to content
// jays.website / business / case study

Business Intelligence — finding what isn't there

Hundreds of thousands of Massachusetts businesses pulled from OpenStreetMap, rolled into comparable zones, and analyzed not for what's present but for what's systematically missing — the underserved markets. Then the proof that the analysis is actionable: the same pipeline productized into live niche directories.

334,638 businesses 885 ZIP zones 14 MA counties OSM-sourced
The problem

Opportunity is an absence, and absences are hard to see

It's easy to count the businesses in a town. It's much harder to answer the question that actually matters to someone deciding where to open: what's missing here that thrives in comparable places? That's not a number you can look up — it only exists by comparing a place against its peers, across many sectors at once. Do it by eye and you'll miss it; do it manually across a whole state and you'll never finish.

I wanted a system that treats a gap as a first-class signal — ranking the places where a given service is conspicuously under-represented relative to similar places.

Approach

Measure every zone against its peers

Pull the full business map, roll it up into comparable zones, and establish what "normal" density and diversity look like for a zone of a given size and type. Then compare each zone against that peer baseline to surface the sectors it's missing, and rank zones by a curiosity score — how systematically under-served they are. The last step is the one that proves it's real: feed the winning markets into a directory factory that spins the analysis into a live product.

Architecture

From map data to market gaps to a product

The analysis answers "where are services missing?"; the directory factory turns that answer into revenue:

Business intelligence pipeline architecture Four tiers. Ingest: OpenStreetMap business points of interest and ZIP/town/county zone definitions feed a PostgreSQL database that geocodes and rolls records up by zone. Compute: per-zone density and diversity metrics, peer-relative sector-gap analysis, and curiosity scoring run in sequence. Deliver: a zone explorer, an interactive spatial map, and a directory factory that productizes selected markets into live niche directories. INGEST STORE COMPUTE DELIVER OpenStreetMap MA business POIs Zone definitions ZIP · town · county PostgreSQL — geocode + zone rollup density + diversity per-zone baseline sector-gap analysis vs. peer zones curiosity scoring rank the absences Zone explorer sector drill-down Interactive map full dataset Directory factory → live products
The intelligence layer picks the market; the directory factory ships the product — the same loop, per vertical.

The full subsystem tour — the density/diversity metrics, the curiosity score, the factory config pattern — is on the business architecture deep-dive.

Key technical decisions

What I chose, and what it cost

Gaps are peer-relative, not absolute

A zone's missing sectors are measured against comparable zones of similar size and type — not against a flat statewide average that would flag every rural town as "missing everything."

Tradeoff Choosing the peer set is a modeling decision that shapes every result, so it has to be defensible, not convenient.

Filter bad geocodes at the door

Coordinate-bounds filtering rejects mis-geocoded records before they reach the map or the stats — a real bug class this pipeline hit and fixed.

Tradeoff A few legitimate edge-of-state records get dropped; cleaner aggregates are worth it.

A directory factory, not a directory

Every vertical-specific detail — brand, copy, services, FAQs — lives in a single config file. Launching a new vertical is: clone template, edit one config, point at a fresh database.

Tradeoff Up-front abstraction cost, repaid the second (and every subsequent) directory.

Isolated database per vertical

Each directory gets its own database, so one vertical's data problem can never take down another.

Tradeoff More databases to operate, in exchange for blast-radius isolation.

Analysis targets the product

The curiosity ranking isn't an academic output — it's the targeting system for which vertical and which cities the factory builds next.

Tradeoff Ties research priorities to commercial ones; the upside is that the analysis has to earn its keep.

Static-fast delivery

Directory sites and the explorer are static and cheap to run, matching the economics of the rest of the suite.

Tradeoff Content freshness is batch, not live — fine for directory data that changes slowly.
Outcomes

From a statewide map to a shipped product

334,638
businesses mapped
src: summary.json
885
ZIP-level zones
src: spatial_summary.json
10
sectors classified
src: summary.json
14
MA counties
src: MassGIS / state
  • Every business, classified. 334,638 places (333,486 in MA) sorted into ten sectors — Professional (51,905), Food & Drink (47,370), Retail (38,335), Healthcare (32,674), and more.
  • Three zoom levels. Rolled up to 885 ZIP zones, 1,028 towns, and county level, so a gap can be read at whatever altitude the decision needs.
  • Gaps ranked, not just shown. Each zone carries a curiosity score and its top missing industries — a ranked answer to "where would a new service face the least competition?"
  • Proven in market. The pipeline's playbook shipped as the Powerwash Directory — a live niche directory (1,445 companies across 20 city pages) built on Next.js + hosted Postgres.

Honest note on counts: the architecture deep-dive still cites an early sample (≈9,862 businesses / 274 zones). The current data export covers 334,638 businesses across 885 ZIP and 1,028 town zones — I've used those here and flagged the deep-dive page for update. The Powerwash Directory counts are as reported by that project and worth re-verifying at merge.

What I'd do next

The honest backlog

  • Automate vertical selection. The curiosity ranking should directly nominate the next directory vertical and its launch cities, instead of a human reading the map.
  • Enrich the gap signal. Business presence alone is one dimension; layering demographics and spending power would separate "under-served" from "can't-support-it."
  • Refresh cadence. OSM data drifts; a scheduled re-extract would keep the gaps current the way the trading and rental pipelines already self-update.
  • Reconcile the deep-dive numbers. Update the architecture page off the live export so the whole site tells one consistent story.

Explore the zones live, or read the deeper engineering tour.