Industry Education
What Is a CRE Knowledge Graph, and Why Does Tenant Rep Need One?
Parcel numbers, business UBIs, permits, and lease records. Every piece of Seattle tenant intelligence is public, and none of it shares a join column.
By Casey Krueger, Founder & CEO, BrokerHQ · Published August 16, 2026 · 11 min read
What is a knowledge graph, in plain terms?
A spreadsheet stores things. A knowledge graph stores things and how they connect.
In a spreadsheet, a building is a row and a tenant is a row in a different file, and the connection between them lives in a column somewhere that you hope was filled in correctly. In a graph, the building is a node, the company is a node, and "occupies suite 1800, lease expires 2027-09-30" is an edge between them. So is "is a subsidiary of." So is "was represented by." So is "is located 400 feet from."
The practical difference is what you can ask. Cherre, which has done more published engineering writing on CRE knowledge graphs than anyone else, put it this way in 2019: "Knowledge graphs are only as good as the questions that drive them. For example, we use knowledge graphs of available CRE data to answer questions like: Who is the true owner for the property? Which properties has this owner bought and sold in the past five years? Which lenders are seeing larger than average number of defaults?"
Notice that none of those questions can be answered by looking up one record. Each one requires walking from a node to a node to a node. Cherre's CTO Ron Bekkerman made the point again in 2021: "A knowledge graph is particularly useful for revealing hidden relationships between entities, by traversing the graph from one node to another over the edges."
Zillow built the same thing for residential search and described the same motivation, which was to "understand and enrich structured and unstructured data from a variety of sources and normalize them to the same vocabulary." Zillow says the graph is what let it ship natural-language search in real estate.
Did the big CRE brokerages build knowledge graphs?
This is where the marketing and the engineering diverge, and it is worth being precise because the answer is not what you would assume.
JLL announced Falcon in October 2024 as an AI platform "featuring multi-modal AI foundation models, data pipelines, security and privacy features, natural language and semantic processing layers and advanced analytics capabilities," and reported that more than 47,000 JLL professionals had used JLL GPT. CBRE's Chief Digital and Technology Officer Sandeep Davé described Ellis AI in February 2025 as "the industry's first self-service multi-model GenAI platform," capable of "querying structured and unstructured data."
Read both announcements closely and neither one says knowledge graph. We looked. Cushman & Wakefield describes a data lake. Moody's CRE describes automated ingestion of rent rolls and cash flow statements. CoStar's public positioning is field research and property records, with 39 years of site inspections behind it.
So the accurate description of the current state is this: the largest brokerages built conversational interfaces on top of their own proprietary data. The companies that built actual graphs are the data-infrastructure players, Cherre most notably, which was acquired by RealPage this year and claims to resolve more than four billion entities, and the residential portals. That is a vendor claim and unaudited, but the published engineering behind it is real.
None of them built one for Seattle office tenant movement. That is not a knock on any of those firms. It is a description of where the gap sits.
Why does search alone not solve this?
Because vector search retrieves text that sounds like your question, and it does not traverse relationships.
Microsoft Research put the limitation plainly when it published GraphRAG in February 2024: "Baseline RAG struggles to connect the dots. This happens when answering a question requires traversing disparate pieces of information through their shared attributes in order to provide new synthesized insights." The peer-reviewed version of that work, published on arXiv in April 2024 and revised in February 2025, opens with the same finding: "RAG fails on global questions directed at an entire text corpus."
Translate that into a tenant-rep question. "Which Seattle companies that grew headcount this year are in buildings where the landlord has a maturing loan and whose leases roll before Q3 2027" is four hops. Company to headcount signal, company to occupancy, occupancy to building, building to owner, owner to debt. There is no document anywhere that contains that answer. A search index cannot retrieve a document that does not exist. A graph can assemble the answer by walking edges.
The honest version of this claim, which we will get to in the counter-argument, is not that graphs beat search. It is that relationship-shaped questions need a relationship-shaped index.
What actually makes this hard? Identity, not storage.
Here is the problem stated as concretely as we can state it.
Take one Seattle office building. In King County's assessor file it is a parcel number. In the Recorder's Office it appears in deeds and liens under an ownership entity that may be a single-purpose LLC with a name that tells you nothing. On the street it has an address that will be written at least three different ways across three databases. It has a building name that exists in no registry anywhere. And it carries a different proprietary identifier inside every vendor system that tracks it.
Now take one tenant in that building. It has a Washington state UBI, the Unified Business Identifier issued by the Department of Revenue, which is searchable by business name, trade name, owners and officers, address, and county. It has a Seattle business license tax certificate with a legal name, a trade name that is usually different, a main location address, and a NAICS code. It may have pulled a tenant-improvement permit through SDCI under a contractor's name entirely. And the name on the lease may be a subsidiary of the name on the door.
Nothing in that paragraph shares a key with anything in the previous one. The parcel file joins on APN. The license file joins on UBI and a mailing address. The permit file joins on a street address. There is no column that connects them.
This is the entity resolution problem, and it is genuinely hard rather than tedious. Cherre described the same mess: "addresses and names (people and corporations) can come in different formats with spelling variations or typos. There's also a need to disambiguate common names." Placekey, the open standard for physical place identity launched in October 2020 by SafeGraph with founding partners including Esri, CARTO, Regrid, and Cherre, published a blunt assessment of address matching in particular: "Getting a 100% match on addresses would be a pipe dream. Even 80% would be considered successful. But in most cases, it's typically 70%." Placekey sells the fix, so treat that number as a vendor estimate, but the direction is corroborated by independent geocoding research: one comparative study across 49 states found match rates of 98%, 82%, 81%, and 30% depending on which vendor processed the same input addresses.
Placekey also names the specific failure that matters most in commercial office. "A single address can have multiple delivery points. Think of individual suites in a very tall skyscraper." A 30-story tower is one address, one parcel, and dozens of tenant entities. Address-level matching cannot see inside it.
And the fragmentation is structural, not local. Regrid, which aggregates parcel data nationally, describes the situation this way: "Every county in the country maintains its own parcel database in its own format, with its own quirks and field names."
What public data exists in Seattle, and what are the rules?
More than most brokers realize, in fragments, with a real legal constraint attached.
The pieces are these. King County Assessor publishes parcel and assessment data for download, with parcel-level lookup through eReal Property and geographic files through King County GIS Open Data. The Recorder's Office carries deeds, easements, mortgages, liens, and excise tax records. The City of Seattle publishes an Active Business License Tax Certificate dataset with legal name, trade name, main location address, expiration date, and NAICS code. SDCI publishes issued building permits and land use permits. The Washington Department of Revenue business lookup covers all current active accounts plus five years of closed accounts, searchable by UBI.
The UBI is the most useful under-used identifier on that list. It is a stable, state-issued company key that survives a DBA change, which is exactly the property you need when the name on the lease and the name on the door do not match.
Now the constraint, and this is the part most CRE AI content ignores entirely. To download King County assessor data you must accept a statement that reads, in part: "Washington State law, RCW 42.56.070(9), prohibits the use of lists of individuals for 'commercial purposes'… 'commercial purposes' means that the person requesting the record intends that the list will be used for general business purposes, including but not limited to communicating with the individual(s) named in the record for the purpose of facilitating profit-expecting activity."
Read that carefully before building anything. The data is public. Turning a list of individuals from it into a prospecting call list is restricted. Building market intelligence about buildings and business entities is a different activity than extracting individuals to cold-call, and any system built on Washington public records has to hold that line deliberately rather than accidentally. We are not lawyers and this is not legal advice; get your own counsel before you operationalize anything here.
What would a CRE knowledge graph actually change for a tenant rep?
Three things, in order of how soon they matter.
The first is that a question becomes answerable at all. Not faster, answerable. "Show me every company in the Denny Regrade under 30,000 square feet whose lease rolls in the next six quarters and that pulled a permit in the last year" is not a search. It is a traversal. Today the way a broker answers it is by pulling a CoStar list, cross-checking it against memory, and calling people.
The second is that the answer carries its provenance. Microsoft's GraphRAG work noted that graph-grounded generation "provides the provenance, or source grounding information, as it generates each response," which lets a person "quickly and accurately audit the LLM's output directly against the original source material." For an audience that a 2026 survey found trusts AI for research and lease abstraction but not for deal decisions, the audit trail is not a nice-to-have. It is the entire adoption question.
The third is compounding. A resolved entity stays resolved. Every time you connect a UBI to a parcel to a suite, the next question that touches any of those three nodes gets cheaper to answer. A list does not do that. A list goes stale the day you export it.
The honest counter-argument
Four objections, and the third one is the strongest.
Graphs are not automatically better. This is the objection we take most seriously because the research supports it. A June 2025 paper out of Hong Kong Polytechnic and Tencent's Youtu Lab opened by noting that "recent studies report that GraphRAG frequently underperforms vanilla RAG on many real-world tasks." Their own benchmark found the accuracy differences modest, with GPT-4o-mini alone scoring 70.68, keyword retrieval 71.66, Microsoft's GraphRAG 72.50, and two graph methods actually scoring below the baseline model. The larger gain showed up in reasoning quality rather than raw accuracy. Their conclusion is the fair one: "GraphRAG's impact varies by question types: it yields significant gains on some types but offers limited benefit for others." Graphs help on multi-hop relationship questions. On simple lookups they add cost and can add noise.
Entity resolution is expensive and never finishes. Bekkerman spelled out the combinatorics: "In a large-scale knowledge graph of, say, 10 million nodes, there would be 50 trillion node pairs, which would make the computation prohibitively slow (that is, many years to run, no exaggeration)." And a partially resolved graph is worse than useless, because "the graph consists of small, disconnected subgraphs, which do not allow any traversing, and therefore makes the graph useless." Half a graph is not half the value.
A broker who knows the market beats a graph. A good Seattle tenant rep already knows which companies are growing and which landlords are stretched, because they talk to people. That knowledge is better than any public-records inference because it includes intent, and intent is not in a permit file. The realistic claim for a graph is that it covers the accounts a broker is not currently thinking about, and that it survives the broker leaving the firm. Neither of those replaces the relationship.
Public records are stale and messy. King County's own disclaimers say the information is subject to change without notice with no warranty as to accuracy, completeness, or timeliness. Permits lag the decision that produced them. Business licenses record a legal address that may be an accountant's office. Any signal derived from public records is probabilistic and should be presented to a broker as a lead worth checking, never as a fact.
Our take
We are building this, so weigh the following accordingly.
The interesting finding from researching this piece was not that knowledge graphs are powerful, which is well established. It was that nobody has built one for tenant movement in a specific office market, and the reason is not technical difficulty. Cherre proved the approach works in CRE. Zillow proved it works in real estate search. Placekey exists because the industry already agreed that place identity is the bottleneck.
The reason it has not been built for tenant rep is the same reason there is no good tenant-rep CRM: the revenue in CRE data sits on the supply side, and a graph of tenant demand serves the side of the table with less money in it. That is an economic explanation, not an engineering one, which means it is the kind of gap a smaller company can actually work in.
If you are a Seattle tenant rep, the useful takeaway is narrower than the technology. The next tool worth your attention is not the one with the best chat interface. It is the one that can tell you where its answer came from, and can still answer when the question requires connecting three things that live in three different places. Our methodology page shows how we source and date every data point we publish.
Sources
Third-party sources
- Cherre, "Building a Knowledge Graph Using Messy Real Estate Data" (2019) and "Improving Knowledge Graph Quality with Neighborhood-Based Entity Resolution" (2021) (graph definition, entity-resolution difficulty, node-pair combinatorics), third-party, vendor engineering blog
- RealPage acquisition announcement of Cherre (four billion entities resolved), third-party, vendor claim and unaudited
- Zillow Tech Hub, "Leveraging Knowledge Graphs in Real Estate Search" (normalizing sources to one vocabulary, natural-language search), third-party, vendor-published
- JLL, Falcon press release, October 29, 2024 (platform description, 47,000 professionals using JLL GPT), third-party, vendor PR
- CBRE, "Everyone Can Have a (Digital) Assistant, with Ellis AI," February 13, 2025 (Ellis AI description, Sandeep Davé quote), third-party, vendor-published
- Microsoft Research, "GraphRAG: Unlocking LLM discovery on narrative private data," February 13, 2024 (baseline RAG failure mode, provenance in generated answers), third-party
- Edge et al., "From Local to Global: A Graph RAG Approach to Query-Focused Summarization," arXiv:2404.16130 (global-question failure), third-party, peer-reviewed preprint
- GraphRAG-Bench, arXiv:2506.02404, and "When to use Graphs in RAG," arXiv:2506.05690 (accuracy table and the finding that GraphRAG often underperforms vanilla RAG), third-party, used as counter-evidence
- Placekey address standardization documentation and placekey.io (70% typical match rate, multiple delivery points per address), third-party, vendor-favorable estimate
- "Geographic disparity of positional errors and matching rate of residential addresses among geocoding solutions," Annals of GIS (98%, 82%, 81% and 30% match rates across 49 states), third-party, peer-reviewed, residential rather than CRE
- Regrid, nationwide parcel documentation (per-county format fragmentation), third-party, vendor-published
- King County Assessor data download terms quoting RCW 42.56.070(9), plus King County mapping disclaimers on accuracy and timeliness, third-party, primary government source
- City of Seattle Open Data, Active Business License Tax Certificate dataset, and SDCI issued building and land use permits, third-party, primary government source
- Washington Department of Revenue Business Lookup (active accounts plus five years of closed accounts, UBI search), third-party, primary government source
BrokerHQ measurements
- BrokerHQ reading of the full JLL Falcon and CBRE Ellis AI announcements, August 16, 2026 (neither announcement uses the term knowledge graph), BrokerHQ measurement, observation about the announcements rather than their internal architecture
- BrokerHQ review of CRE data-vendor positioning, August 2026, covering Cushman & Wakefield, Moody's CRE and CoStar (data lake, rent-roll ingestion and field-research positioning), BrokerHQ measurement, attributed estimate
Disclosure: BrokerHQ is building tenant-rep data infrastructure for Seattle office brokers and competes with several of the products named above. This analysis was AI-assisted using BrokerHQ's research corpus and reviewed by Casey Krueger. It is not legal advice.
Liked this?
The weekly Seattle CRE Brief brings the same kind of read to your inbox every Friday. Subscribe, it's free.