AI & Brokerage Workflow

    The Audit Trail Is the Liability: What AI Legal-Research Errors Mean for Tenant-Rep Brokers

    Stanford tested three AI legal research tools against vendor no-hallucination claims and found error rates of 17-33%. Here's what that means for AI use in tenant-rep deal analysis.

    By Casey Krueger, Founder & CEO, BrokerHQ · Published September 8, 2026 · 8 min read

    17-33%hallucination rate of leading AI legal research toolsSTANFORD REGLAB, PREREGISTERED EVALUATIONAI & BROKERAGE WORKFLOW · SINCE JUN 2026BROKERHQ

    What did Stanford actually find when it tested AI legal research tools?

    Stanford RegLab ran a preregistered evaluation, meaning the test design was locked before the results came in, against three commercial AI legal research products: Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI. All three are marketed as eliminating hallucination risk. The study found each tool hallucinated between 17% and 33% of the time. Accuracy came in at 65% for Lexis+ AI, 41-42% for Westlaw AI-Assisted Research, and 19% for Ask Practical Law AI (Stanford RegLab, arXiv 2405.20362).

    Worth flagging directly: the paper's authors disclose that one of them advises language-model companies, and Stanford said it would augment the study after some parties questioned the results. That's a live caveat, not a settled number. But even a generous discount off a 19-65% accuracy spread doesn't get you to a place where an unverified AI output belongs in a signed lease negotiation.

    Why does this matter for CRE brokers instead of lawyers?

    Because the failure mode isn't legal-specific. It's structural. A tool marketed as reliable that is actually unreliable creates a documentation problem before it creates an accuracy problem. If a broker uses an AI tool to pull a rent comp, model an escalation clause, or flag a lease-expiry date, and the tool is wrong 20-30% of the time while claiming near-zero error, the broker now has two exposures instead of one: the underlying mistake, and the fact that they trusted a system whose own accuracy claims don't hold up.

    One Seattle-adjacent broker on Reddit put it plainly: telling a client "the AI said so" makes the negligence case stronger, not weaker, because "no idiot would leave a contractual obligation to AI without manually verifying when there are dollars on the table" (r/CommercialRealEstate, cited in BrokerHQ's own sourced data). A second broker on the same thread reached the identical conclusion independently: "no sane person leaves a $1M obligation to a black-box AI. Verification is the only way to kill negligence." Two brokers, two separate posts, same structural read. That's not one anxious outlier, that's a pattern.

    How common is this concern among brokers right now?

    BrokerHQ's own sourced data tracks AI liability risk as an active complaint theme: 9 logged complaints as of September 7, 2026, ranked 8th out of 102 tracked themes by frequency, first observed June 15, 2026. The pattern shows up across 2 independent platforms, meaning it's not confined to a single forum's echo chamber. Intensity is tagged high across every quote tied to the theme.

    What's not yet knowable from this data: whether frequency is rising, flat, or falling. The measurement window doesn't have four weeks of trailing history stored yet, so there's no trend line, only a current snapshot. Treat the 9-complaint count as a real, sourced number and the absence of a trend as an honest gap, not a hidden one.

    What does 'auditable' actually require in a tenant-rep deal?

    A third broker, discussing AI use in CRE modeling more broadly, framed the real risk correctly: "The biggest risk I'd watch is less 'can AI build the model?' and more 'can someone else audit the assumptions quickly?'" (r/CommercialRealEstate, cited in BrokerHQ's own sourced data). That's the operative distinction. A model that's fast and internally consistent but opaque to a second reviewer, a lender, an equity partner, a client's attorney, doesn't reduce your liability. It just moves the liability into a black box with your name on the transmittal.

    Auditability means someone other than the broker who ran the tool can trace an output back to its inputs in minutes, not hours. If a rent comp, an escalation calc, or a lease-abstraction result can't be walked back to a verifiable source in front of a client, it's not ready to go in a deal memo.

    Isn't this just AI-skepticism dressed up as a liability argument?

    Fair challenge. The honest version of the counter-case: hallucination rates in general-purpose legal AI tools don't automatically transfer to narrower CRE use cases like pulling a public rent comp or summarizing a lease clause, where the task is more constrained and the ground truth is easier to check. It's also true that human brokers make errors too, and no one's holding manual comp-pulling to a 100% accuracy bar. The Stanford numbers are specific to legal research tools tested against legal-research tasks; extrapolating a 19-33% hallucination rate onto every AI use case in CRE is a stretch the study itself doesn't make.

    What doesn't survive the counter-case: the marketing gap. The Stanford study's most damaging finding wasn't the error rate, it was that vendors claimed "hallucination-free" performance while the tools hallucinated on a quarter of queries. That gap between claim and behavior is the liability, independent of whether the underlying task is narrow or broad. A broker who can show they verified an AI output manually has a defense. A broker who can show only that a vendor promised accuracy has a plaintiff's exhibit.

    BrokerHQ's View

    BrokerHQ's view: this isn't an argument against using AI in tenant-rep work. It's an argument against using it without a receipt. If a broker can't show, in under five minutes, exactly where a rent comp or lease-abstraction number came from, that broker has taken on a liability the vendor's marketing copy didn't disclose. The fix isn't avoiding AI tools, it's refusing to let any AI output into a client-facing deliverable without a traceable source attached. That's the standard BrokerHQ builds to, and it's the standard we'd tell any broker to hold a vendor to before trusting a black-box number with a client's signature on the line.

    FAQ

    What did the Stanford study find about AI legal research tools?

    Stanford RegLab's preregistered evaluation found Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate between 17% and 33% of the time, despite vendor marketing claiming hallucination-free performance. Measured accuracy was 65%, 41-42%, and 19% respectively (Stanford RegLab, arXiv 2405.20362).

    Does telling a client 'the AI said so' protect a broker from liability?

    No. Two brokers in BrokerHQ's own sourced data independently argued the opposite: citing an unverified AI output to a client strengthens a negligence case, because it demonstrates no one manually checked a high-dollar obligation before acting on it.

    How common is AI liability concern among brokers right now?

    BrokerHQ's own sourced data logs 9 AI liability risk complaints as of September 7, 2026, ranking it 8th of 102 tracked complaint themes, observed across 2 independent platforms since June 15, 2026. There isn't yet enough measurement history to say whether this is rising or falling.

    What makes an AI-generated deal analysis 'auditable'?

    A broker in BrokerHQ's own sourced data framed it as whether someone else can quickly audit the assumptions behind the model, not just whether the model runs. If a rent comp or lease calculation can't be traced back to its source in minutes, it isn't ready for a client-facing deliverable.

    Does this mean brokers should avoid AI tools entirely?

    The evidence doesn't support that conclusion. The documented risk is the gap between vendor accuracy claims and actual tool behavior, not AI use itself. The operative fix is requiring a traceable, verifiable source behind every AI-assisted number before it reaches a client.

    Sources

    BrokerHQ measurement

    • BrokerHQ's own sourced data on AI liability risk: 9 complaints in the current tracking period, high intensity, corroborated on 2 independent forums/platforms, tracked since 2026-06-15, counts measured as of 2026-09-07 (BrokerHQ measurement, not a published third-party source)

    Third-party

    • Stanford RegLab, arXiv 2405.20362: preregistered evaluation finding Lexis+ AI, Westlaw AI-Assisted Research, and Ask Practical Law AI hallucinate 17-33% of the time, with measured accuracy of 65%, 41-42%, and 19% (study published 2024, cited in BrokerHQ research August 11, 2026) (third-party)
    • r/CommercialRealEstate, two brokers independently arguing that citing an unverified AI output to a client strengthens a negligence case rather than reducing it (third-party forum, anecdote)
    • r/CommercialRealEstate, broker framing the real AI risk in CRE modeling as whether someone else can quickly audit the assumptions behind the model (third-party forum, anecdote)

    Disclosure: This analysis was AI-assisted using BrokerHQ's proprietary research corpus.

    Liked this?

    The weekly Seattle CRE Brief brings the same kind of read to your inbox every Friday. Subscribe, it's free.