AI & Brokerage Workflow

    The AI Liability Gap No One's Underwriting For

    Stanford found legal AI tools hallucinate 17-33% of the time. Seattle tenant-rep brokers are running lease analysis through AI with no verification layer. Here's the exposure.

    By Casey Krueger, Founder & CEO, BrokerHQ · Published September 1, 2026 · 7 min read

    17-33%hallucination rate of leading legal AI toolsSTANFORD REGLAB, PREREGISTERED EVALUATIONAI & BROKERAGE WORKFLOW · SINCE JUN 2026BROKERHQ

    How often does legal AI actually hallucinate?

    Stanford RegLab ran a preregistered evaluation of the three biggest names in AI-assisted legal research: Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI. All three were marketed as hallucination-free or as having eliminated the hallucination problem. The study found each tool hallucinated between 17% and 33% of the time, with accuracy ranging from 65% down to 19% depending on the product (Stanford RegLab, arXiv 2405.20362). Worth flagging: one of the paper's authors discloses an advisory relationship with language-model companies, and Stanford said it would augment the study after some parties questioned the results. That doesn't erase the finding. It means treat the exact numbers as directional, not gospel, and the direction is bad enough on its own.

    That's legal research, not lease analysis. But the tools brokers are pulling into underwriting and lease abstraction workflows are built on the same class of model, doing the same kind of pattern-matching over dense contractual language. There's no reason to assume tenant-rep lease review is immune to the same failure mode. There's also no published study telling us it's better. That absence is the point: nobody's measured it, and brokers are shipping decisions off it anyway.

    Why does the gap matter more than the error rate?

    One line from the research buried in an unrelated audit says it cleanest: 'the damage came from the gap, not the error rate.' A 20% hallucination rate is survivable if every output gets checked before it touches a client deliverable. It becomes a liability event the moment nobody checks, because the failure isn't visible until the lender, the landlord, or opposing counsel finds it first.

    That's the exact shape of what's showing up in BrokerHQ's own sourced data right now. This theme, AI liability exposure in underwriting and lease analysis, is sitting at high intensity across two independent broker sources since mid-June. Zero blog posts cover it. Zero citation probes have tested for it. It's a demand-with-no-supply gap, and it's not a small one: this is the kind of topic that gets cited by name when a broker or their E&O carrier goes looking for precedent after something goes wrong.

    What are brokers actually saying about this?

    The broker voice on this is sharper than any vendor whitepaper. From r/CommercialRealEstate: 'no sane person leaves a $1M obligation to a black-box AI. Verification is the only way to kill negligence.' Same thread, a harder version of the same point: 'But AI said so makes an even STRONGER case for negligence because no idiot would leave a contractual obligation to AI without manually verifying when there are dollars on the table.'

    That's not a broker worried about AI being wrong. That's a broker worried about what a plaintiff's attorney does with the phrase 'the AI told me to.' It converts a good-faith tool choice into evidence of reduced diligence, which is the opposite of what E&O coverage assumes a broker is doing.

    A second broker, on a different thread, names the actual audit problem: 'The biggest risk I'd watch is less can AI build the model, and more can someone else audit the assumptions quickly. If the model is custom to how one operator thinks, that's useful internally, but lenders and equity partners may still care about standardization and traceability.' That's the whole gap in one sentence. A model that works for you isn't the same thing as a model someone else can check.

    What does a verification workflow actually need to include?

    Known: Stanford's numbers give a hard floor for what 'unverified' error rate looks like in an adjacent high-stakes domain. Broker quotes establish that the exposure is understood at the practitioner level, not just a compliance abstraction. Thought but unconfirmed: whether lease-analysis AI tools specifically perform better or worse than the legal-research tools Stanford tested, because nobody's run that study yet. Untested: whether documented human verification actually reduces E&O claims in practice for firms using AI-assisted underwriting, because there isn't enough case history yet to know.

    Given what's known, the one move that's defensible right now: treat every AI-assisted lease abstraction or underwriting output as a draft that requires a named, dated, human sign-off before it leaves the building. Not because it guarantees the AI was right. Because it's the difference between an error and a negligence claim. The broker quote said it best: the AI-said-so defense doesn't reduce liability, it increases it, because it's an admission nobody checked.

    Isn't a verification step just slower, and doesn't that defeat the point of using AI at all?

    Yes, it's slower. That's the trade, and it's worth naming directly instead of waving away. If the whole pitch for AI-assisted underwriting is speed, adding a mandatory human check claws back a meaningful chunk of that speed gain. For a broker running high volume on low-stakes deals, that math might genuinely not be worth it. But the evidence here is specifically about high-stakes exposure: a $1M obligation, a lender relationship, a client deliverable with your name on it. On those deals, the speed you gained isn't worth what you lose if the unverified number is wrong and there's no paper trail showing anyone looked. The counter-argument holds for the low-stakes tail of the business. It doesn't hold for the deals big enough to be the subject of a broker complaint in the first place.

    BrokerHQ's View

    BrokerHQ's view: the Stanford number is a legal-research finding, not a lease-analysis finding, and I'm not going to pretend otherwise. But the gap it exposes is the same gap showing up in our own sourced data, at high intensity, with nobody writing about it and nobody testing for it. That combination, real adjacent-domain error data plus zero coverage, is exactly the kind of thing worth publishing on before someone's E&O carrier makes them learn it the hard way. The fix I'd actually recommend, in order: put a named human sign-off on every AI-assisted number that touches a client or lender deliverable, before you touch anything about model selection or vendor claims. Reasoning: the broker quotes in this data aren't arguing about which AI tool is more accurate, they're arguing about who's accountable when it's wrong. Accountability is a workflow problem, not a model problem, and it's the one you can fix this week without waiting on a vendor to publish better numbers.

    FAQ

    What hallucination rate did Stanford find in AI legal research tools?

    Stanford RegLab's preregistered evaluation found Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI each hallucinated between 17% and 33% of the time, with accuracy scores of 65%, 41-42%, and 19% respectively, despite vendor marketing claiming hallucination-free results.

    Does this hallucination data apply directly to AI lease analysis tools?

    Not directly. The Stanford study tested legal research tools, not lease-analysis or underwriting AI. There's no published study measuring hallucination rates specifically in tenant-rep lease review workflows. That gap in the research is itself worth noting: brokers are adopting these tools without a domain-specific error rate to calibrate against.

    Why do brokers say using AI increases negligence exposure rather than reducing it?

    Because if a lease analysis or underwriting number turns out wrong and the record shows no human verified it, the AI said so explanation reads as an admission that due diligence was skipped, not as a mitigating factor. As one broker put it, that framing makes the negligence case stronger, since no reasonable professional would leave a contractual obligation this size to an unverified model.

    What's the single highest-priority fix for firms using AI in underwriting?

    A documented, named human sign-off on every AI-assisted output before it reaches a client or lender. This addresses the accountability gap directly, which is what the broker complaints and the Stanford gap-versus-error-rate finding both point to, and it doesn't require waiting on better AI accuracy numbers to implement.

    Is this AI liability topic currently covered anywhere in tenant-rep content?

    No. BrokerHQ's sourced data shows this theme at high intensity across two independent broker sources since mid-June 2026, with zero blog content and zero citation probes addressing it, making it a demand-with-no-supply gap in the category.

    Sources

    • BrokerHQ sourced data on AI liability exposure: 8 complaints, high intensity, corroborated on 2 independent sources, first seen 2026-06-15, counts measured as of 2026-08-31 (BrokerHQ measurement, not a published third-party source)
    • Stanford RegLab, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools" (arXiv 2405.20362): 17% to 33% hallucination rates and accuracy of 65%, 41-42%, and 19% across Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI (third-party)
    • Reddit r/CommercialRealEstate, broker comments on negligence exposure and the AI-said-so defense (third-party forum, anecdote)
    • Reddit r/CommercialRealEstate, broker comment on whether someone else can audit a custom model's assumptions quickly (third-party forum, anecdote)

    Disclosure: This analysis was AI-assisted using BrokerHQ's proprietary research corpus.

    Liked this?

    The weekly Seattle CRE Brief brings the same kind of read to your inbox every Friday. Subscribe, it's free.