Business Intelligence — finding what isn't there
Hundreds of thousands of Massachusetts businesses pulled from OpenStreetMap, rolled into comparable zones, and analyzed not for what's present but for what's systematically missing — the underserved markets. Then the proof that the analysis is actionable: the same pipeline productized into live niche directories.
Opportunity is an absence, and absences are hard to see
It's easy to count the businesses in a town. It's much harder to answer the question that actually matters to someone deciding where to open: what's missing here that thrives in comparable places? That's not a number you can look up — it only exists by comparing a place against its peers, across many sectors at once. Do it by eye and you'll miss it; do it manually across a whole state and you'll never finish.
I wanted a system that treats a gap as a first-class signal — ranking the places where a given service is conspicuously under-represented relative to similar places.
Measure every zone against its peers
Pull the full business map, roll it up into comparable zones, and establish what "normal" density and diversity look like for a zone of a given size and type. Then compare each zone against that peer baseline to surface the sectors it's missing, and rank zones by a curiosity score — how systematically under-served they are. The last step is the one that proves it's real: feed the winning markets into a directory factory that spins the analysis into a live product.
From map data to market gaps to a product
The analysis answers "where are services missing?"; the directory factory turns that answer into revenue:
The full subsystem tour — the density/diversity metrics, the curiosity score, the factory config pattern — is on the business architecture deep-dive.
What I chose, and what it cost
Gaps are peer-relative, not absolute
A zone's missing sectors are measured against comparable zones of similar size and type — not against a flat statewide average that would flag every rural town as "missing everything."
Tradeoff Choosing the peer set is a modeling decision that shapes every result, so it has to be defensible, not convenient.Filter bad geocodes at the door
Coordinate-bounds filtering rejects mis-geocoded records before they reach the map or the stats — a real bug class this pipeline hit and fixed.
Tradeoff A few legitimate edge-of-state records get dropped; cleaner aggregates are worth it.A directory factory, not a directory
Every vertical-specific detail — brand, copy, services, FAQs — lives in a single config file. Launching a new vertical is: clone template, edit one config, point at a fresh database.
Tradeoff Up-front abstraction cost, repaid the second (and every subsequent) directory.Isolated database per vertical
Each directory gets its own database, so one vertical's data problem can never take down another.
Tradeoff More databases to operate, in exchange for blast-radius isolation.Analysis targets the product
The curiosity ranking isn't an academic output — it's the targeting system for which vertical and which cities the factory builds next.
Tradeoff Ties research priorities to commercial ones; the upside is that the analysis has to earn its keep.Static-fast delivery
Directory sites and the explorer are static and cheap to run, matching the economics of the rest of the suite.
Tradeoff Content freshness is batch, not live — fine for directory data that changes slowly.From a statewide map to a shipped product
- Every business, classified. 334,638 places (333,486 in MA) sorted into ten sectors — Professional (51,905), Food & Drink (47,370), Retail (38,335), Healthcare (32,674), and more.
- Three zoom levels. Rolled up to 885 ZIP zones, 1,028 towns, and county level, so a gap can be read at whatever altitude the decision needs.
- Gaps ranked, not just shown. Each zone carries a curiosity score and its top missing industries — a ranked answer to "where would a new service face the least competition?"
- Proven in market. The pipeline's playbook shipped as the Powerwash Directory — a live niche directory (1,445 companies across 20 city pages) built on Next.js + hosted Postgres.
Honest note on counts: the architecture deep-dive still cites an early sample (≈9,862 businesses / 274 zones). The current data export covers 334,638 businesses across 885 ZIP and 1,028 town zones — I've used those here and flagged the deep-dive page for update. The Powerwash Directory counts are as reported by that project and worth re-verifying at merge.
The honest backlog
- Automate vertical selection. The curiosity ranking should directly nominate the next directory vertical and its launch cities, instead of a human reading the map.
- Enrich the gap signal. Business presence alone is one dimension; layering demographics and spending power would separate "under-served" from "can't-support-it."
- Refresh cadence. OSM data drifts; a scheduled re-extract would keep the gaps current the way the trading and rental pipelines already self-update.
- Reconcile the deep-dive numbers. Update the architecture page off the live export so the whole site tells one consistent story.
Explore the zones live, or read the deeper engineering tour.