Skip to content
// jays.website / rental / case study

Rental Intelligence — a state as a shortlist

Every multifamily building in Massachusetts assembled into one spatial database, scored 0–100 against real HUD rent benchmarks, and connected to an ownership graph of thirty thousand landlords. Not a listings site — a decision engine that turns a state's worth of parcels into a ranked set of acquisition targets.

305,476 buildings $290B assessed value 14 counties 30,251 owners mapped
The problem

The targets are public — and impossible to see

Everything you need to find a good multifamily acquisition is technically public: parcel geometry, assessed values, unit counts, owner names, market rents. It's just scattered across town assessor databases, a statewide GIS layer, and federal rent tables that no one thinks to join together. So the work becomes manual cross-referencing, one town at a time — slow, inconsistent, and blind to the owner-level patterns that actually signal a deal.

I wanted to collapse that into a single question: across the entire state, which buildings are the best opportunities right now, and who owns them?

Approach

One spatial database, one deal score, one ownership graph

Join the authoritative sources into a single spatial schema, score every building on the same 0–100 scale, and resolve hundreds of thousands of parcel records down to the actual people and entities that own them. The score blends estimated yield against the HUD benchmark with distress and ownership signals — and, critically, a data-confidence component so a building with thin assessor data can't masquerade as a strong signal. The engine produces a shortlist, not a verdict; the final call stays human.

Architecture

From 305,476 parcels to a ranked map

A heavy database does the thinking at home; the public site is just the static export of its conclusions:

Rental intelligence pipeline architecture Four tiers. Ingest: MassGIS statewide parcels, town assessor records, and HUD Fair Market Rents feed a PostgreSQL/PostGIS spatial database. Compute: a 0–100 composite deal score and an owner-name-resolution graph read the database. Deliver: results become county opportunity rankings, an interactive statewide map, and a daily static JSON export committed to git. INGEST STORE COMPUTE DELIVER MassGIS parcels statewide geometry Assessor records units · value · year HUD Fair Market rent benchmarks PostgreSQL + PostGIS — one spatial schema composite 0–100 deal score ownership graph name resolution County rankings top opportunities Interactive map statewide, street-level Static JSON export daily auto-commit
Heavy DB at home → static export at the edge. A daily refresh re-scores and auto-commits the data.

The subsystem-level tour — data hygiene, the scoring components, the ownership resolution — is on the rental architecture deep-dive.

Key technical decisions

What I chose, and what it cost

Heavy database, light site

PostgreSQL/PostGIS does every spatial join and rollup; the public site reads only a pre-exported static dataset. The database never sits in the request path.

Tradeoff Rankings are as fresh as the last daily export, not live — accepted in exchange for a site that's fast, cheap, and can't be knocked over by traffic.

Confidence is part of the score

Each building's 0–100 score carries a data-quality component, so a sparse assessor record can't produce a confident ranking.

Tradeoff Some genuinely good buildings with thin data rank lower than they deserve — a deliberate bias toward not crying wolf.

Two-layer ranking

The engine separates "good building" from "good opportunity right now," so quality and timing don't get blended into one muddy number.

Tradeoff More scoring logic to maintain, but the shortlist answers the question an investor actually asks.

Resolve owners, not parcels

Owner-name resolution collapses parcel records into distinct owners, surfacing portfolio holders and absentee landlords — because who owns a building is itself a signal.

Tradeoff Name resolution is fuzzy and imperfect; it's treated as a strong hint, never ground truth.

Shortlist, not verdict

The system ranks; a human decides. A review-labeling loop feeds real accept/reject calls back toward the weights.

Tradeoff It won't make the decision for you — by design, because the cost of a wrong acquisition dwarfs the convenience.

Enforce schema keys everywhere

Early mismatched keys silently broke joins across services. Keys are now enforced consistently — a hard lesson in how quietly a spatial join can rot.

Tradeoff More rigid contracts between services, which is exactly the point.
Outcomes

A whole state, queryable in seconds

305,476
buildings in the database
src: investment_summary.json
$289.7B
aggregate assessed value
src: investment_summary.json
30,251
owners resolved
src: ownership graph
14
counties covered
src: county_breakdown
  • Statewide, in one place. Every MA county — from Middlesex (59,521 buildings) to Nantucket — scored on the same scale and queryable from one map.
  • 14,982 live opportunities. Buildings currently surfaced with an active opportunity score, re-ranked on every daily refresh.
  • Owner-level intelligence. Portfolio holders and absentee landlords are flagged, turning a parcel list into a set of people to actually call.
  • Self-updating. A scheduled daily refresh re-scores the dataset and commits the export — every day's rankings are reproducible from git.

Honest note on counts: earlier pages cited 339,856 buildings and "177k scored." The live pipeline export now reports 305,476 buildings and 14,982 carrying an active opportunity score — so I've switched to those and flagged the difference. The gap looks definitional (all buildings stored vs. those passing the opportunity threshold); it's worth pinning down.

What I'd do next

The honest backlog

  • Nail down the "scored" definition. The 305k-stored vs 15k-scored gap should be one documented number with a clear threshold — ambiguity in a headline metric is a credibility leak.
  • Close the feedback loop. The accept/reject review labels exist; wiring them back into the scoring weights as a real learning signal is the next lever.
  • Rent estimation where listings are thin. HUD FMR is a floor, not a market rent; a per-building rent model would sharpen yield estimates in hot submarkets.
  • Ship the Pro tier. Alerts, CSV export, and a cap-rate calculator are scoped — the natural first paid surface on top of a free explorer.

Explore the rankings live, or read the deeper engineering tour.