Watchline NYC — Accountability infrastructure for NYC housing enforcement
Frequently asked questions

How Watchline works — and what to trust

Straight answers for the two rooms that ask the hardest questions: people steeped in SQL and public housing data, and people rightly skeptical of AI.
The one-line answer to everything: the substance is the public record; the graph and the language model are tooling around it, and every element is labeled sourced or inferred — a lead to verify, never a legal determination.

The project

What is Watchline, and how is it different from Who Owns What?

Watchline makes NYC's public housing record legible — who owns and runs the city's housing — conversationally, and with its uncertainty labeled. It builds on JustFix's Who Owns What (WoW), which it credits and complements, not replaces. Its main addition is an owner-identity layer that resolves one owner's differently-named LLCs into a single owner, and a separation of three questions the record often conflates: who manages a building, what operational network it runs through, and who actually owns it.

What data does it use?

Only NYC public records — HPD registrations, complaints, violations and vacate orders; DOB violations; ECB judgments; ACRIS deeds and mortgages; and marshal evictions — assembled into a knowledge graph. No private or proprietary data.

How current is the data?

It's a periodic snapshot, not a live feed, and each published page is dated. Figures can lag the live public record — re-check against the source before acting on anything.

The methods

You could do all this in SQL — what does the knowledge graph add?

Most of the pipeline is SQL: WoW is Postgres, the record linkage runs in DuckDB, and every conditions aggregate is a GROUP BY. The graph earns its place for the parts that are transitive and multi-relationship: connected-component owner identity, walking several hops across owner / manager / address / deed links, and graph data science (centrality, community detection, link prediction). "Count violations per owner" is SQL; "everyone within three hops across four kinds of link, ranked by structural importance" is the graph.

There's AI in this — should I trust it?

Be skeptical; that's the right instinct, and the design assumes it. Two different things get called "AI" here, and only one does the substantive work. The model that decides which records are the same owner is an auditable statistical record-linkage model — not a neural network — and its precision is measured. The language model is only a front door: it turns a plain-English question into a read-only query and reports what the query returned; it never invents an answer. Every element comes back labeled sourced or inferred. If you don't trust the language model, ignore it and read the sourced records it points you to — the findings don't depend on it.

How do you decide two differently-named LLCs are the same owner?

A high-confidence "same owner" link (CONNECTED_BY_SPLINK) connects records that resolve to one owner. It comes from three sources: a name-anchored probabilistic record-linkage model (precision-first, with vetoes for common names and shared aggregator offices), hand-curated overrides for cases the model can't reach, and exact same-registered-LLC matches. It never links two different surnames. It's what collapses one owner's typo'd offices and shell LLCs into a single entity.

How does your "portfolio" relate to Who Owns What's?

The operational-nexus layer (:Portfolio) is WoW's construction, reproduced — the same clustering over WoW's own name/address connections — plus the "same owner" links above. Because links only add, it only ever merges what WoW split; it never fragments WoW's groupings. Separately, the owner-identity layer (:OwnerGroup) is stricter: it uses only the identity links and drops the shared-address glue, and it's the layer behind statements like "these 115 LLC names are one owner."

A concrete case: landlord Ramon Escobar's single Bronx office (2432 Grand Concourse #504) is entered in the registration data a dozen inconsistent ways — GRAND CONCOURSE, GRAND COURSE, GRAND COCNOURSE, the apartment dropped, differing city and ZIP. Because Who Owns What links on the exact address string, each variant peels buildings off into a separate portfolio (WoW: 24 + 2). The sharpest culprit isn't even the misspelling: Who Owns What requires an exact ZIP match on every link, and two of these buildings ended up with a blank ZIP in standardization — so even the one spelled perfectly gets split off. Watchline reunites all 26 because it keys on resolved owner identity and the name-free deed, never the raw string (see the case study).

Trust & limits

How accurate is it?

On a 105-record hand-adjudicated gold set, the owner resolution is essentially never wrong when it links two records — precision ≈ 1.0, zero cross-surname merges. What is not yet done is the head-to-head accuracy comparison against Who Owns What where the two disagree: a blind evaluation of ~530 sampled pairs is built but not run. The two systems diverge on tens of thousands of buildings, in both directions; which is right where they disagree is exactly what that evaluation will decide.

Isn't naming people a privacy concern?

Everything shown is public record, and tools like Who Owns What already publish these names with caveats. Watchline presents the same records — often more accurately, for example by separating owners that address-based grouping wrongly lumps together. Every ownership link is an inference, labeled as such — a lead to investigate, never a legal determination — and the public examples are pinned to figures already adjudicated in the public record.

Using it

Can I use it? Is it a product? Is it open source?

It's an independent sabbatical prototype, not a shared service. It reads public data only. The point of publishing these examples is to gather feedback and collaborators — especially to help run the blind accuracy evaluation and sanity-check the three-layer model. The code lives at github.com/bobflagg/WatchlineNYC.