Legal work is becoming one of AI’s most aggressively pursued markets, but the pitch arrives with a fairly large asterisk shaped like a court citation. Google has put Gemini Enterprise for Legal into preview, Anthropic has released Claude Legal Solutions, and SpaceXAI has created a legal-solutions page for Grok. The common promise is not simply faster drafting. It is a safer workflow: models that can work with legal databases, retrieve authorities and help professionals investigate documents without inventing the law.

That is an important shift in emphasis. Generative AI’s worst legal failures have generally not come from a lawyer using software to organize a large document set or locate a case. They have come when an attorney treats a fluent paragraph as completed research and files it without checking whether the cited decision exists, says what the filing claims, or applies to the issue at hand.

The consequences have already been substantial. The Ontario Law Society Tribunal ordered lawyer Shahryar Mazaheri to pay $31,150 in costs in June after filings drafted with an earlier version of Grok included fabricated citations and invented legal principles. The tribunal had found in December that the tool had been used to prepare a factum, meaning a written legal argument, for an appeal. Its assessment of the resulting material was blunt: it was described as “gibberish.”

The episode is not an argument that lawyers can never use AI. It is a warning that a tool’s ability to generate persuasive-sounding language is fundamentally different from its ability to establish legal authority. That distinction is now at the center of legal AI marketing, and it will remain central to how courts and law firms judge these products.

Why citations are the pressure point

A legal citation is not decoration. It gives the opposing side and the court a route to inspect the authority being relied on. A judge, statute, regulation or prior decision can be checked for its wording, jurisdiction, procedural posture and relevance. If the authority is fictional, the argument loses its foundation immediately. If it is real but mischaracterized, the result can be nearly as damaging.

Large language models, or LLMs, generate responses by predicting likely language patterns. That can make them adept at summarizing material placed before them or shaping a draft into coherent prose. But when asked to supply a citation from memory, a model can combine believable court names, dates, reporters and legal concepts into something that looks authentic without corresponding to an actual case. This is commonly called a hallucination: an output presented as fact that is false or unsupported.

The Ontario tribunal stressed a related problem in its costs decision: an LLM does not exercise judgment, grasp nuance or use a moral compass, and it is inclined to return an answer rather than acknowledge uncertainty. In legal practice, that last tendency is especially dangerous. A research tool that cannot find a directly relevant case should leave a gap for the lawyer to investigate. A text generator that fills the gap with a plausible invention can turn incomplete research into a serious professional failure.

Mazaheri’s sanction was not imposed merely because AI was involved. The failure was the lack of thorough verification before the material was filed. That distinction matters because it establishes the practical rule for every new legal AI offering: using a product does not move the attorney’s responsibility onto the product’s developer.

The evidence that this is not an isolated error

Fabricated legal citations have appeared in work from individual lawyers, self-represented litigants and major firms. Damien Charlotin’s database has cataloged nearly 2,000 cases worldwide involving more than 800 lawyers and 1,100 people representing themselves. In Canadian courts, reported cases involving fake citations increased from seven in 2024 to 86 in 2025, with another 39 reported in the first quarter of 2026.

Even a large, highly resourced firm is not immune. Sullivan & Cromwell submitted an emergency letter to the Southern District of New York in April after a bankruptcy motion filed with Chief Judge Martin Glenn was found to contain 42 AI-generated fabrications. The incident cuts through the comforting assumption that such mistakes occur only when inexperienced users experiment with consumer chatbots. The core risk is procedural: an unverified output passed through a workflow and into a filing.

There is also a useful historical marker in the abandoned effort by DoNotPay to have an AI “robot lawyer” argue a court case in early 2023. State bar associations threatened criminal prosecution, and the plan was dropped. That moment concerned the unauthorized practice of law and courtroom representation. Today’s products are more commonly positioned as tools for legal professionals rather than replacements for them, but the human professional remains the accountable actor.

Google and Anthropic are responding to the citation problem with a product design often described as grounding. In this context, grounding means tying a model’s answer to supplied or connected information rather than asking it to rely solely on patterns learned during training. A grounded system may retrieve material from a connected database, use that material as the basis for an answer, and show the user where the supporting information came from.

Gemini Enterprise for Legal is being previewed with Gottlieb, Freshfields, Weil and Williams & Connolly. Its design routes questions through connectors to external legal databases, including Everlaw and NetDocuments. Google Cloud CEO Thomas Kurian has framed factual accuracy and grounding in legal authority as critical to these agentic workflows.

An agentic workflow refers to a system that can carry out multiple steps toward a task, rather than only produce a one-off answer. In legal use, that could mean taking a request, searching connected repositories, assembling relevant information and preparing an initial work product for review. The phrase does not itself guarantee accuracy. It describes a working process, not a finding that every result is correct.

Anthropic’s Claude Legal Solutions takes a comparable connector-led approach. The product reportedly includes 20 connectors to legal platforms and 12 pre-built plugins. Anthropic has reported a 90.9% result for Opus 4.7 on the BigLaw Bench legal reasoning benchmark, and Freshfields and Quinn Emanuel are among its customers.

Benchmarks can be useful signals, but they answer narrower questions than lawyers face in practice. A benchmark score can indicate performance on a defined collection of tasks. It does not, by itself, prove that every citation in a real filing is complete, current, correctly quoted and applicable in the relevant jurisdiction. Nor does it establish that a professional can skip the normal task of reading the cited authority.

Retrieval is a safeguard, not a substitute for judgment

The practical advantage of a connected legal tool is clear. Instead of telling a model to invent a case citation from its internal training, a lawyer can have it retrieve information from a legal research platform or document system. The source can then be checked directly. That makes the process more auditable and may sharply reduce a familiar category of hallucination.

But retrieval solves a limited set of problems. It may confirm that a case exists, while still leaving difficult questions: Is it binding or merely persuasive? Has it been overturned, narrowed or distinguished? Does the quotation omit a qualification? Is a fact pattern materially different? These are questions of legal analysis, not just information lookup.

They also explain why an answer with links or citations should not automatically be treated as verified. A traceable citation gives the reviewer a path. Verification is the act of taking that path, reading the authority and confirming that the proposition fairly follows from it.

That is where the ordinary legal workflow should remain deliberately unglamorous. A lawyer using AI-generated research or drafting should inspect each authority, confirm the relevant text, consider its legal status and ensure the final argument accurately reflects the record and governing law. The time saved in gathering or sorting information cannot be confused with permission to stop reviewing it.

Independent testing remains a missing piece

The new products are entering a market with reason for caution. A 2024 Stanford RegLab study tested purpose-built legal AI platforms from LexisNexis and Westlaw and found hallucination rates of 17% and 33%. Those figures concern different tools from Google’s and Anthropic’s latest offerings, so they are not a direct measurement of Gemini Enterprise for Legal or Claude Legal Solutions. Still, they demonstrate why a legal branding label is not enough to establish reliability.

There has not yet been a comparable independent audit of Google’s or Anthropic’s legal tools. Until one exists, customers have vendor descriptions, benchmark claims and the particulars of each deployment, but not an equivalent independent measure of how these systems behave across the messy range of real legal tasks.

That gap is especially notable because legal work can fail in ways that are subtle. An obvious fabricated citation is bad, but a genuine case used for the wrong proposition can be harder to catch and potentially more misleading. Evaluations that measure only whether a response sounds useful may miss the very errors that matter most in a filing.

SpaceXAI’s Grok-based legal offering presents an even less clear picture from the available details. A Cursor blog post described Grok 4.5 as appropriate for finance, legal work and other computer-based tasks. But SpaceXAI’s later announcement focused on coding, agentic tasks and knowledge work without specifically discussing legal uses. Its legal page does not explain an architecture or a way to trace citations to primary sources. SpaceX officially acquired Cursor for $60 billion on August 15, placing that legal marketing under the Grok umbrella, but a dedicated page is not the same thing as a disclosed verification system.

What responsible adoption looks like

The legal AI debate can become needlessly binary: either these systems are forbidden because they can err, or they are treated as automated experts because they can produce useful work quickly. The evidence points toward a more demanding middle ground. Tools with access to verified external sources have a better starting position than tools asked to recall law from a model’s training. Yet the attorney filing the document must still validate the result.

  • Use connected, authoritative materials where possible. Retrieval from legal databases or firm repositories is more defensible than requesting uncited legal propositions from a general model.
  • Open every important authority. Confirm the decision exists and supports the precise statement for which it is cited.
  • Review context, not only quotations. A sentence can be accurate in isolation and misleading in a different factual or procedural setting.
  • Treat benchmarks as limited evidence. A score reflects a specified evaluation, not a blanket assurance that a tool is safe for every matter.
  • Keep accountability human. Bar rules place the responsibility for filed citations on the attorney, regardless of whether AI assisted with the work.

This is the same broader governance issue facing software that takes action or supplies high-stakes recommendations. Systems can add approvals and long-term context, as explored in the debate around always-on agents and approval controls, but controls matter only when a responsible person actually uses them.

For the legal profession, the decisive question is therefore not whether AI can write an argument that looks lawyerly. It plainly can. The question is whether the process allows a lawyer to establish that each legal claim is real, relevant and appropriately used before it reaches a court. The newest products are attempting to make that task easier through retrieval and connectors. They cannot make it optional.