Pantea opened her own law firm this year. On a run together a few weeks ago, she described line items she had not planned for.
Clients take a contract she has drafted, run it through a chatbot, and send back the results to her. The feedback always has errors. Some of the feedback covers language already handled. Some misread what the law requires. Others are flatly wrong about the legal standard.
She has answers to all of it, in writing, at her hourly rate.
The clients had gone looking for a free second opinion. It showed up on their invoice anyway.
Despite using AI daily herself for years, “sloppy AI” use by others has increased Pantea’s net work load. Clients and pro se claimants – those who represent themselves without the benefit of an attorney – are now sending verbose memos that say very little. Buyers of her clients’ services send AI readouts of the contract leaving her to sort through it. She routinely sends the reports back asking for redlines. The redlines are rarely better.
Even when attorneys are involved, it is usually obvious when AI is used because they submit non-relevant markups and take extreme positions on virtually all contractual matters.
Someone has to review and respond.
Start with what’s working
Legal services fail to reach many
The client evaluating a legal document through AI was doing the sensible thing. They were paying for a document they couldn’t fully evaluate, in a domain where not understanding something gets expensive later. A tool exists that will explain it in plain language, instantly, for free. The same can be said about someone without a lawyer running a contract through AI for comments.
Most people never get near a lawyer. Law is a highly specialized and sub-specialized field with an average hourly rate of $349 per hour. In a survey of litigators, more than three quarters said their firms turn away cases that are not cost-effective, with $100,000 case value being the most commonly named cutoff. Below that line, the practical answer the profession has been giving people for decades is you’re on your own.
The Legal Services Corporation’s national survey found that low-income households received no legal help, or not enough, for 92 percent of the civil legal problems that substantially affected them. Three in four had at least one such problem in the previous year. And in the best behavioral study we have, only 8 percent of people’s civil justice claims reached a lawyer and 8 percent reached a court. Asked why, 56 percent called it bad luck, part of life, or God’s plan. So when 14 percent of consumers told Clio they’d used AI for a legal question, and another 43 percent said they would, they weren’t defecting from lawyers. Most of them were never going to have one anyway.
People are turning to AI instead. In OpenAI’s study of how people use ChatGPT, non-work messages grew from 53 percent of consumer traffic in mid-2024 to more than 70 percent a year later. About 29 percent of all conversations are people asking for practical guidance. In the BCG field experiment run by Fabrizio Dell’Acqua and colleagues, consultants using GPT-4 inside the model’s competence finished more work, faster, at higher quality. The bottom half of performers gained more than twice as much as the top half.
Judges are seeing it too. Maritza Braswell, a federal magistrate in Colorado, says she’s getting better-drafted pleadings from self-represented litigants and can generally understand their arguments better with AI assistance than without it.
This is not an argument that people should stop using AI. It’s an argument about what happens next.
Two costs, but only one is reduced by AI
A claim has two prices: producing it and checking whether it’s true.
AI collapsed the first. A plausible, well-organized, citation-dense objection to a contract clause now costs about eleven seconds and nothing more.
AI has done nothing for the person who has to check the output. Checking a filing still requires somebody who knows the law in that jurisdiction, reads the clause, knows the rest of the document, and understands how the two interact. That takes exactly as long as it always did. The sheer volume of AI output makes it take longer.
Alberto Brandolini named this phenomenon more than a decade ago: refuting nonsense takes an order of magnitude more energy than producing it. With AI, we have industrialized production and left refutation to human beings.
Economists have started modeling it. In a February 2026 paper, Catalini, Hui, and Wu argued the binding constraint on the economy is no longer intelligence. It is “human verification bandwidth”: the capacity to validate, audit, and take responsibility when execution becomes abundant. They give the behavior a name that fits Pantea’s inbox: unverified deployment becomes privately rational, a Trojan Horse externality. Sending the output costs you nothing. Checking it costs someone else a great deal.
This happens every day mostly out of sight. Researchers at BetterUp Labs and Stanford surveyed 1,150 US employees about workslop: AI output that looks like work and doesn’t advance the task. Forty percent had received some in the previous month, and each instance cost the recipient close to two hours. Nobody billed anybody. It came out of people’s evenings.
Law is different. Outside the courts, the hours are already being counted. The additional time to sort through AI slop doesn’t disappear in a lawyer’s practice. It gets itemized in the next invoice.
Why the almost-right answer is the expensive one
The costly AI output is not the output that’s badly wrong. It’s the output that’s almost right.
Charles Hatley, a managing partner who now regularly meets clients showing up with fully drafted proposals from ChatGPT, described the pattern in one line: AI legal work is “rarely obviously incorrect.” It’s plausible, it cites things, and it blends the statutes of three states into an argument that would be sound in none of them.
Judge Braswell’s full quote contains an important qualification. She’s seeing better-drafted pleadings, and she has to be careful, because some of them contain hallucinations and errors.
Both halves are true. The second is expensive. An incoherent filing gets dismissed in a paragraph. A filing that is articulate, formally correct, and wrong in one buried respect has to be read closely by someone qualified to spot it.
Raising the floor on the writing raised the cost of the review.
The same BCG experiment shows the flip side. When Dell’Acqua’s team set a task designed to fall just outside the model’s competence, the group working without AI got it right 84.5 percent of the time. The groups with GPT-4 managed 60 and 70 percent. The treatment group that had received prompt-engineering training did worse.
Dahl and colleagues at Stanford measured hallucination rates of 58 to 88 percent on direct, verifiable questions about federal cases. The rates got worse in lower courts, in smaller jurisdictions, and on harder questions. That is a description of the person who could not get a lawyer.
What changed in the person asking
The study that reframed this for us isn’t about law.
Marcoccia, Quattrociocchi, and Capraro ran five experiments with 3,132 participants. Four were preregistered, meaning the researchers filed their predictions publicly before they collected any data, so they could not change the question after seeing how it came out. People were asked hard questions about film trivia and could always decline to answer. The questions were picked because the AI reliably got them wrong, which separates using AI from AI being right.
Without AI, people answered “I don’t know” on 44% of questions. With AI, that fell to 3%. AI-using participants answered far more and were correct about a third as often. The participant’s reported confidence went from 30 to 76 on a hundred-point scale. Researchers even tested paying people for right answers and fining them for wrong ones with moderate success. But it did not bring them anywhere near as accurate as people working without AI. Up to that point, people chose whether to see the AI’s answer. In the last experiment the researchers showed it to everyone, unasked. People still stopped saying “I don’t know.” The task was trivia, so this is not evidence that AI makes people worse at everything.
This is not a knowledge problem, it’s a metacognitive one. The question was never whether someone knows the answer. It’s whether people can tell when they don’t know it and when they stop asking questions about if they have the correct answer. Recognition of uncertainty is what keeps people from acting on a wrong answer. Twenty years of studying how people decide things at work has convinced Shonna that it is also the most fragile. It goes first under time pressure, first under social pressure, and now, apparently, first when something fluent is willing to answer on your behalf.
That is the mechanism behind the invoice.
Before, a client who didn’t understand section 7 sent a question. “I don’t follow this part, can you walk me through it?” A question is cheap. Answering it takes four minutes and often strengthens the relationship.
Now the same client sends a verdict. “Section 7 is unenforceable.” A verdict has to be disproved. It takes research, a written response, and the delicate work of telling a paying customer that the confident thing they were told is not correct. The uncertainty didn’t get resolved. It got converted into an assertion, and assertions are expensive.
Further, the client cannot tell which of the fifteen objections in the AI output is worth raising. That is the whole reason they used the tool. The expertise you lack when you read the contract is the same expertise you lack when you read the machine’s opinion of the contract. This means you cannot price your own question before you ask it. You find out what it was worth after you have paid to find out.
Across two experiments, 698 people used ChatGPT to work through logical reasoning problems from the law school admission test. Their scores came out three points above the test’s norms. Their estimates of their own scores came out four points above what they had actually earned, and the people who rated themselves most AI-literate were the furthest off.
The more destructive AI mistake: telling you that you’re right
The discussion so far has been about the model being wrong. A second failure type leaves no trace at all.
Fabricated citations are detectable. A court can find it, a database can count it, a judge can sanction it. The data on AI being wrong comes from this. But a model agreeing with you and telling you that you have a good case produces no artifact. Nothing to check, nothing to sanction, and it happens before the filing.
In March, Science published three preregistered experiments by Cheng and colleagues involving 2,405 participants across eleven frontier models. The models affirmed users’ actions 49 percent more often than humans did “even when queries involved deception, illegality, or other harms.” In the third study, participants brought a conflict from their own lives. After one conversation, their conviction that they were in the right rose sharply and their willingness to repair the situation fell.
The mechanism the models used were even more interesting. Sycophantic responses were significantly less likely to mention the other person, or that person’s point of view. The model doesn’t argue you into being right. It narrows the frame until only your side is in it.
A related benchmark found that models decline to challenge a user’s framing 88 percent of the time, against 60 percent for human responders, and let ungrounded assumptions stand in 86 percent of cases. Researchers took a single conflict, wrote it up from each person’s point of view, and gave both versions to the models. In 48 percent of cases, the model told both people they were in the right. A separate 2026 benchmark asked whether models notice when a legal question is missing the facts that would decide it. Researchers stripped material facts out of real questions, and mostly the models answered confidently without flagging the gap. When the missing fact was where the case stood, whether a deadline had passed or a court had already ruled, they caught it about one time in ten.
This is what the legal professional was for. The very first lecture many law students receive on their first day of class is about professional judgement. Herbert Kritzer’s survey of contingency-fee practice found that roughly two out of every three people who contact a plaintiff’s lawyer are turned away, and the most common reason given is not that the case is too small. It’s that there’s no liability. At one large firm’s actual intake, 45 percent of callers were rejected on the first phone call and only about 15 percent ended up represented. A lawyer working on contingency stakes unpaid hours on your case having merit. That makes “you don’t have a case” the cheapest sentence they can say, and often the most useful one you will hear.
Now read the AI labs’ own documentation. Anthropic’s most recent system card reports that the model is “actively worse on the broader ‘wet blanket’ metric for dismissive or discouraging output,” and adds that this “is potentially linked to its improvement on sycophancy.” Discouragement is a tracked defect. The most useful sentence a lawyer can say to you is “you don’t have a case.” The labs treat that as a mistake.
To our knowledge, nobody has run the study that connects validation to filing. The Science result measures willingness to apologize, not willingness to sue. The researchers who documented the filing surge attribute it to cheaper drafting rather than to changed beliefs about the merits. Cheap drafting is the less troubling explanation. The worse one is that AI is convincing people their case is stronger than it is, and they file believing they were wronged. They are the hardest kind to deter, because sanctions, fee-shifting, and dismissal all assume someone who knows better.
One case later in this piece began exactly that way. When a plaintiff’s attorney told her the settlement could not be reopened, she asked a chatbot whether she was being gaslighted. It told her she was.
The unwilling person in the room
Everything so far involves people who chose to be there. The client hired the lawyer. Pantea took the client.
Now consider someone who chose nothing.
Federal pro se civil filings averaged about 23,000 a year from 2005 through 2022. In fiscal 2025 they hit 41,490. Self-represented plaintiffs account for 59 percent of the growth in non-prisoner federal civil filings, and filings by represented parties stayed flat, meaning new people are entering the system pro se. In randomly sampled federal civil complaints, AI-detected text went from a 0.1 percent baseline before 2023 to 18 percent by early 2026.
Most of these suits fail, and the failure rate has not changed. Self-represented plaintiffs win well under 1 percent of the cases they file. Among the few that reach a judgment on the merits, they prevail 4 percent of the time, against 51 percent when both sides have lawyers. The researchers who assembled this data state the asymmetry plainly: a federal agency sued by a plaintiff with AI-drafted filings “faces an adversary whose marginal filing cost has fallen, while the agency’s own response cost has not.”
But a lawsuit doesn’t have to succeed to cost the person on the other end. In the median federal civil case with any discovery, a defendant spends about $20,000 in 2008 dollars, roughly $31,000 today, and that figure excludes every case resolved in under 60 days. Even the cheapest possible exit isn’t free: in one 2021 fee ruling a federal court reviewed the invoices for a single motion to dismiss and found 67.6 hours and $21,437 in legal fees.
Consider a case that ran the course. A woman settled a disability benefits dispute in January 2024, with a full release and dismissal with prejudice. In 2025 she fed her attorneys’ communications to ChatGPT, which per the pleadings told her she’d been “gaslighted,” that the settlement was “invalidated,” and encouraged her to fire her lawyers. It drafted her arguments to reopen the case. The motion was denied. She filed a second suit the same day. Across the two cases she generated more than 60 filings across two cases, citing at least one case that does not exist.
The insurance company alleges roughly $300,000 in defense costs.
To try to recover it, the company filed a third lawsuit, in March 2026, against OpenAI.
That is not a quirk of strategy. Under the American Rule, each side in US litigation pays its own attorneys regardless of who wins. The defendant’s only route to being made whole ran through a novel tort theory against a technology company, because the ordinary route does not exist.
The courts were built to turn people away
The American court system is designed not to hear every claim. Bringing a case means meeting deadlines, paying fees, following procedural rules written for specialists, setting out your facts in writing, and making legal arguments you have to research yourself. Then come motions, discovery, and the rules of evidence. Every step is a filter.
Cost and complexity are the barriers people name. Time belongs on the list too. Somewhere in the fifth hour of drafting a claim and working out where to file it, a person can stop being angry enough to keep going. The rules never called that a cooling-off period, but it worked as one.
Society made that trade intentionally. We accepted that people with valid claims would not be heard, in exchange for keeping out claims that weren’t worth bringing and keeping the judiciary affordable to run. You can argue about whether it was the right trade. One side of it has now changed, and the other has not been adjusted to match. Courts are not ready for the influx.
An aggrieved person can produce something that looks like a legal claim in minutes with AI. Federal pro se civil filings rose, and the number of federal judges did not. Courts were already slow, with motions waiting years on crowded dockets, and cases with genuine merit will now wait behind cases that would once have died in the fifth hour.
Defense lawyers bill by the hour, so the influx pays them, and they report the strain anyway. Their clients are out of pocket.
Why the obvious fixes don’t fix it
The instinct is to say courts should punish someone for filing something frivolous. But the bar for court sanctions is purposefully high and the rules for doing it were written for lawyers. They rarely touch anyone else.
In April 2026, the federal Advisory Committee on Civil Rules took up whether to write a rule about AI-hallucinated filings. It voted unanimously to drop the question. The reasoning was that Rule 11 already covers it.
In the US database, 59 percent of these cases in America are pro se litigants. Every rule they pointed to was written for lawyers. Rule 11 only works on someone who will be back, has money, and has a license to lose. Federal Rule of Appellate Procedure 46 governs attorney discipline. Bar referral, malpractice exposure, and the ABA’s ethics opinion on generative AI do not apply to an unrepresented person.
In the same meeting, the Committee also declined a proposal that would have required courts to help self-represented litigants. It called the idea worthy and said the Civil Rules gave it no authority to act on it.
So in one sitting it declined to tighten the duty and declined to build the help that would let people meet it. The Fifth Circuit, which had drafted a rule that would have covered self-represented filers, abandoned it after commenters said existing rules already reached counsel. Ten months later the same court wrote: “It is a problem that is getting worse—not better.”
Federal law contains one limit on filing volume, and it works by taking away the fee waiver. After three dismissals for frivolousness, a prisoner loses the right to file without paying. Anyone who can pay the filing fee faces no volume limit at all.
The only other tool is a case-by-case injunction, which a court can issue after a pattern has already been established, and which the Ninth Circuit calls “an extreme remedy that should rarely be used.” The one limit that counts filings was aimed at the population least able to produce them.
What works, and what we’re doing to it
There is a version of this problem we’ve already solved once.
Legislatures long ago concluded that for some suits the filing itself is the injury, and built machinery to match. For example, Anti-SLAPP statutes are designed to deter meritless, retaliatory lawsuits that silence free speech by punishing the act of filing. Currently 40 states and DC have Anti-SLAPP statutes on the books which require: dismissal before discovery, mandatory fees to the prevailing defendant, an automatic stay of discovery, and an immediate right of appeal if the case is not dismissed.
Texas went further and authorized sanctions sized not only to compensate the defendant but “to deter the party who brought the legal action from bringing similar actions.” California maintains a public registry of roughly 3,600 people barred from filing without permission, in continuous publication since 1991, and the statute lists them by how much they file rather than by any single injury they caused.
The design has one idea in it. Get the defendant out before discovery starts, and make the person who filed pay for the meritless filing. Punishment afterward was tried in other corners of the law and mostly failed because of how hard collection is after the fact.
Which is what makes the last eighteen months strange to watch.
The mechanism around early screening filing requirements is being chipped away. In October 2025 the Ninth Circuit eliminated immediate appeal from denied anti-SLAPP motions. In January 2026 the Supreme Court overruled a state law barrier to filing medical malpractice claims in federal courts by holding in Berk v. Choy that heightened state filing requirements do not apply to federal filings. With no federal Anti-SLAPP laws, courts were already split about whether Anti-SLAPP rules applied in federal courts, and after the Supreme Court decision in Berk commentators concluded that they do not. In June the Supreme Court declined to revisit any of it.
During the same eighteen months where courts have been shedding early-exit screenings, AI-generated text in federal complaints went from about 1 percent to 18 percent. Case loads were already at a breaking point. According to the Administrative Office of the U.S. Courts, it takes 27 months from filing for a civil case to go to trial.
One change on each side of the invoice
If you’re the one holding the AI output, send the question, not the verdict.
Instead of “AI flagged section 7 as unenforceable,” write: “I ran this by a chatbot and it raised something about section 7 that I don’t understand. Can you tell me whether there’s anything there?”
Identical information. A fraction of the cost, because it puts the burden back on explanation rather than refutation, and it restores the sentence the research says we’re losing. You don’t have to stop using AI to prepare for a conversation with an expert. You have to stop letting it reach your conclusions for you.
If you’re the professional, quote the question before you answer it.
The instinct is to price this work as a line item. But the hours are already the client’s, and they pay for them either way. The client already has the price. What they lack is the judgment to know whether the spend is worth it, which is the same judgment they hired you for.
So supply it before you do the work. Read the list, and tell them: working through all fifteen points will take about three hours, two of them look substantive to me, and the rest are covered elsewhere in the document. Which would you like? That is a four-minute email, and it gives the client a decision they cannot make alone. It also converts the verdict back into a question, which is where this all started.
If you’re the professional using AI, don’t have it replace your professional judgment. Check its work, remove extreme positions, and use your judgment on whether something should go forward.
And if you work anywhere near how these rules get made, the question on the table is not whether people should use AI to understand their own legal problems. They will, and for many of them the alternative was nothing at all. The question is who checks, and who pays them to do the checking.
We keep asking whether AI makes people more accurate. The better question is whether it makes them more willing to say “I don’t know.”
Think of the last time you got an AI answer in a domain you don’t practice in. What would you have done in 2020? Asked someone who knew, or let it go?
And did having the answer make you more likely to go check it, or less?
References
Advisory Committee on Civil Rules. (2026, April 14). Agenda book (pp. 438–450) and draft minutes. Administrative Office of the U.S. Courts.
Berk v. Choy, No. 24-440 (U.S. Jan. 20, 2026).
Bebchuk, L. A. (1988). Suing solely to extract a settlement offer. Journal of Legal Studies, 17(2), 437–450.
Colbert, K. (2026, February 5). The surge of pro se plaintiffs. Minnesota Journal of Law & Inequality.
Cheng, M., Lee, S. A., Khadpe, P., Yu, Q., Han, K., & Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792), eaec8352. https://doi.org/10.1126/science.aec8352
Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2026). ELEPHANT: Measuring and understanding social sycophancy in LLMs. In *Proceedings of the International Conference on Learning Representations (ICLR 2026)*. https://arxiv.org/abs/2505.13995
Catalini, C., Hui, X., & Wu, L. (2026). Some simple economics of AGI (arXiv:2602.20946). https://arxiv.org/abs/2602.20946
Charlotin, D. (2026). AI hallucination cases [Database]. https://www.damiencharlotin.com/hallucinations/
Chatterji, A., Cunningham, T., Deming, D., Hitzig, Z., Ong, C., Shan, C., & Wadman, K. (2025). How people use ChatGPT (NBER Working Paper No. 34255). https://www.nber.org/papers/w34255
Clio. (2025). 2025 legal trends report. https://www.clio.com/resources/legal-trends/read-online/
Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93. https://doi.org/10.1093/jla/laae003
Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2025). Navigating the jagged technological frontier. Organization Science. https://doi.org/10.1287/orsc.2025.21838
Di Pietro, S., Carns, T. W., & Kelley, P. (1995). Alaska’s English rule: Attorney’s fee shifting in civil cases. Alaska Judicial Council.
Fernandes, D., Villa, S., Nicholls, S., Haavisto, O., Buschek, D., Schmidt, A., Kosch, T., Shen, C., & Welsch, R. (2025). AI makes you smarter but none the wiser: The disconnect between performance and metacognition. Computers in Human Behavior, 175, 108779.
Fletcher v. Experian Information Solutions, Inc., No. 25-20086 (5th Cir. Feb. 18, 2026).
Gopher Media LLC v. Melone, No. 24-2626 (9th Cir. Oct. 9, 2025) (en banc), cert. denied (U.S. June 15, 2026).
Hatley, C. D. (2026, April 27). When the client brings ChatGPT to the consultation. Virginia Lawyers Weekly.
Institute for the Advancement of the American Legal System. (2010). Civil litigation survey of chief legal officers and general counsel belonging to the Association of Corporate Counsel.
Kim, M. (2026, June 4). How courts are coping with AI lawsuits. MIT Technology Review.
Kritzer, H. M. (2002). Seven dogged myths concerning contingency fees. Washington University Law Quarterly, 80(3), 739–794.
Niederhoffer, K., Kellerman, G. R., Lee, A., Liebscher, A., Rapuano, K., & Hancock, J. T. (2025, September 22). AI-generated “workslop” is destroying productivity. Harvard Business Review.
Legal Services Corporation. (2022). The justice gap: The unmet civil legal needs of low-income Americans. https://justicegap.lsc.gov/
Lee, E. G., III, & Willging, T. E. (2009). Federal Judicial Center national, case-based civil rules survey: Preliminary report. Federal Judicial Center.
Marcoccia, C., Quattrociocchi, W., & Capraro, V. (2026). AI advice suppresses people’s willingness to say “I don’t know,” even when the advice is wrong and accuracy is incentivized.
Sandefur, R. L. (2014). Accessing justice in the contemporary USA: Findings from the Community Needs and Services Study. American Bar Foundation.
Reporters Committee for Freedom of the Press. (2026). Anti-SLAPP legal guide. https://www.rcfp.org/anti-slapp-legal-guide/
Shah, A. V., & Levy, J. Y. (2026). Access to justice in the age of AI: Evidence from U.S. federal courts [Working paper]. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6766859
Vincent, S. J., Calloway, D., Yu, F., Bean, A. M., & Seedat, N. (2026). InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries (arXiv:2608.20220). https://arxiv.org/abs/2608.20220



