
Watchline makes NYC's public housing record legible — who owns and runs the city's housing — conversationally, and with its uncertainty labeled. It builds on JustFix's Who Owns What (WoW), which it credits and complements, not replaces. Its main addition is an owner-identity layer that resolves one owner's differently-named LLCs into a single owner, and a separation of three questions the record often conflates: who manages a building, what operational network it runs through, and who actually owns it.
Only NYC public records — HPD registrations, complaints, violations and vacate orders; DOB violations; ECB judgments; ACRIS deeds and mortgages; and marshal evictions — assembled into a knowledge graph. No private or proprietary data.
It's a periodic snapshot, not a live feed, and each published page is dated. Figures can lag the live public record — re-check against the source before acting on anything.
Most of the pipeline is SQL: WoW is Postgres, the record linkage runs in DuckDB, and every
conditions aggregate is a GROUP BY. The graph earns its place for the parts that are
transitive and multi-relationship: connected-component owner identity, walking several hops
across owner / manager / address / deed links, and graph data science (centrality, community
detection, link prediction). "Count violations per owner" is SQL; "everyone within three hops across
four kinds of link, ranked by structural importance" is the graph.
Be skeptical; that's the right instinct, and the design assumes it. Two different things get called "AI" here, and only one does the substantive work. The model that decides which records are the same owner is an auditable statistical record-linkage model — not a neural network — and its precision is measured. The language model is only a front door: it turns a plain-English question into a read-only query and reports what the query returned; it never invents an answer. Every element comes back labeled sourced or inferred. If you don't trust the language model, ignore it and read the sourced records it points you to — the findings don't depend on it.
A high-confidence "same owner" link (CONNECTED_BY_SPLINK) connects records that resolve
to one owner. It comes from three sources: a name-anchored probabilistic record-linkage model
(precision-first, with vetoes for common names and shared aggregator offices), hand-curated
overrides for cases the model can't reach, and exact same-registered-LLC matches. It never links
two different surnames. It's what collapses one owner's typo'd offices and shell LLCs into a single
entity.
The operational-nexus layer (:Portfolio) is WoW's construction, reproduced — the
same clustering over WoW's own name/address connections — plus the "same owner" links above.
Because links only add, it only ever merges what WoW split; it never fragments WoW's
groupings. Separately, the owner-identity layer (:OwnerGroup) is stricter: it uses
only the identity links and drops the shared-address glue, and it's the layer behind statements like
"these 115 LLC names are one owner."
A concrete case: landlord Ramon Escobar's single Bronx office
(2432 Grand Concourse #504) is entered in the registration data a dozen inconsistent ways —
GRAND CONCOURSE, GRAND COURSE, GRAND COCNOURSE, the
apartment dropped, differing city and ZIP. Because Who Owns What links on the exact address
string, each variant peels buildings off into a separate portfolio (WoW: 24 + 2). The sharpest
culprit isn't even the misspelling: Who Owns What requires an exact ZIP match on every link, and two
of these buildings ended up with a blank ZIP in standardization — so even the one spelled perfectly
gets split off. Watchline reunites all 26 because it keys on resolved owner identity and the name-free
deed, never the raw string (see the case study).
On a 105-record hand-adjudicated gold set, the owner resolution is essentially never wrong when it links two records — precision ≈ 1.0, zero cross-surname merges. What is not yet done is the head-to-head accuracy comparison against Who Owns What where the two disagree: a blind evaluation of ~530 sampled pairs is built but not run. The two systems diverge on tens of thousands of buildings, in both directions; which is right where they disagree is exactly what that evaluation will decide.
Everything shown is public record, and tools like Who Owns What already publish these names with caveats. Watchline presents the same records — often more accurately, for example by separating owners that address-based grouping wrongly lumps together. Every ownership link is an inference, labeled as such — a lead to investigate, never a legal determination — and the public examples are pinned to figures already adjudicated in the public record.
It's an independent sabbatical prototype, not a shared service. It reads public data only. The point of publishing these examples is to gather feedback and collaborators — especially to help run the blind accuracy evaluation and sanity-check the three-layer model. The code lives at github.com/bobflagg/WatchlineNYC.