Steven Schwartz sat in a Manhattan courtroom on June 8, 2023, trying to explain how he had cited six cases that did not exist. He had practiced law for three decades. He had never, he said, understood that ChatGPT could invent a judicial opinion out of thin air. He had even asked the chatbot whether the cases were real. It told him yes. He believed it.
I have tried more than two hundred cases. I have watched witnesses lie under oath with a straight face, and I have watched honest people fall apart because they were scared. Schwartz was neither lying nor scared exactly. He was something worse for a lawyer. He was credulous. He took the word of a machine over the discipline of actually pulling the reporter, reading the opinion, and shepardizing the cite. And a federal judge made an example of him that the whole profession is still absorbing.
the affidavit that should have ended a career
The case was Mata v. Avianca, Inc. (S.D.N.Y. 2023). A personal-injury plaintiff, a knee injured by a metal serving cart on an airplane, a routine fight over whether the claim was time-barred under the Montreal Convention. Nothing about the underlying dispute was memorable. What made it a permanent fixture in every CLE deck since is what Schwartz and his colleague Peter LoDuca filed in opposition to a motion to dismiss.
The brief cited Varghese v. China Southern Airlines. It cited Shaboon v. Egyptair, Petersen v. Iran Air, Martinez v. Delta Air Lines, and more. Avianca’s lawyers went looking for these decisions. They could not find them. Neither could the court. They did not exist. ChatGPT had generated them, complete with fake internal quotations and fake citations to fake reporter volumes, and Schwartz had copied them into a federal filing.
Judge P. Kevin Castel did not accept ignorance as an excuse, and he was right not to. On June 22, 2023, he sanctioned Schwartz, LoDuca, and their firm, Levidow, Levidow & Oberman, ordering them to pay $5,000 and to send letters to the real judges who had been named as the supposed authors of the phantom opinions. Read the opinion. Castel was careful. He wrote that there is nothing inherently improper about using a reliable AI tool for assistance. The sanction was not for using the tool. It was for abandoning the gatekeeping function that every lawyer owes the court, and then, when caught, doubling down instead of coming clean.
That last part matters more than the technology. When the fake cases surfaced, the lawyers did not immediately confess error. They produced the supposed opinions, which ChatGPT had also fabricated on request. They kept the fiction alive. The cover-up, as always, was worse than the crime.
I remember thinking, when I first read Mata, that it would be a one-off. A cautionary tale. A single embarrassed lawyer whose name would become a verb. I was wrong. It was the opening witness in a very long trial.
the docket got crowded fast
Within months, the imitators arrived. Some of them were not obscure solos. They were serious people at serious firms.
In the criminal case of Michael Cohen, President Trump’s former lawyer, a motion filed in late 2023 in the Southern District of New York cited three cases that turned out to be fabrications generated by Google Bard. Cohen, it emerged, had passed the cites to his lawyer David Schwartz, who filed them without checking. Judge Jesse Furman flagged the fictions and demanded an explanation. Cohen filed a declaration admitting he had used Bard and had not realized it could hallucinate case law. Furman declined to impose sanctions in the end, but the humiliation was total and public, and the opinion he wrote in March 2024 was a warning shot to every lawyer who thought a famous client made verification optional.
The Second Circuit weighed in too. In Park v. Kim (2d Cir. 2024), the court dealt with an attorney, Jae S. Lee, whose brief cited a nonexistent decision produced by ChatGPT. The panel referred her to the court’s grievance panel and used the opinion to state plainly that citing fake authority is a violation of an attorney’s basic obligations, full stop. When a federal appellate court takes time in a merits opinion to lecture the bar about robot citations, the problem has stopped being a curiosity.
Then came 2025, and the dam broke.
In February 2025, the Wyoming case Wadsworth v. Walmart, Inc. put lawyers from Morgan & Morgan, one of the largest plaintiff-side firms in the country, in the hot seat. Their motions in limine cited eight cases. Most did not exist. The cites came out of an internal AI platform the firm had built, and one of the attorneys, Rudwin Ayala, had trusted it. Judge Kelly Rankin revoked Ayala’s pro hac vice admission and sanctioned the lawyers. The firm itself scrambled. It circulated an internal memo warning its own attorneys that AI could fabricate authority and that filing hallucinated cases could get them fired. When a firm with hundreds of lawyers has to send that email, you understand the scale.
The MyPillow litigation produced another one. In Coomer v. Lindell (D. Colo. 2025), lawyers representing Mike Lindell filed a brief riddled with defective citations, including cases that were misquoted, misattributed, or simply invented. Judge Nina Y. Wang sanctioned attorneys Christopher Kachouroff and Jennifer DeMaster in the summer of 2025, ordering them to pay roughly $3,000 each. Kachouroff admitted the brief had been run through generative AI and had not been properly checked before filing.
There was Kohls v. Ellison (D. Minn. 2025), which had a special flavor of irony. It was a case about a Minnesota deepfake statute, and one side submitted an expert declaration from a Stanford professor, Jeff Hancock, who studies misinformation. His declaration cited academic studies that did not exist, apparently produced by an AI tool. The court excluded the declaration. A misinformation expert undone by AI-generated misinformation in a case about AI-generated misinformation. You could not write it as fiction and be believed.
Big firms were not immune either. In Alabama, lawyers from Butler Snow, a large regional firm, filed court papers in prison-conditions litigation that contained citations to nonexistent cases generated by AI. The court demanded answers in 2025 and the firm had to publicly own the failure. There was also Gauthier v. Goodyear Tire & Rubber Co. (E.D. Tex. 2024), where an attorney was sanctioned and ordered to pay a penalty and attend legal-education training after filing AI-invented authority.
By the middle of 2025, this was no longer a handful of anecdotes. A French researcher, Damien Charlotin, started keeping a public database of court decisions involving AI-hallucinated citations. It swelled past two hundred documented instances and kept climbing, spanning federal and state courts, plaintiff and defense, solo practitioners and AmLaw firms, and even litigants who were lawyers representing themselves. The pattern repeated with numbing regularity. A brief gets filed. Opposing counsel or the judge’s clerk cannot find a cited case. An order to show cause issues. The lawyer confesses to using ChatGPT, Bard, Gemini, Claude, or some in-house tool, and admits nobody read the actual opinions. Sanctions follow.
The dollar amounts stayed modest in most cases, a few thousand here, a few thousand there. But the reputational cost was not modest at all. These opinions are published. They are searchable. A lawyer’s name attached to a fabricated-citation sanction follows him around like a bad tattoo. Clients read them. Judges remember them. Opposing counsel print them out for the next fee fight.
what the standing orders actually say now
Judges did not wait for the rules committees. They moved on their own, and they moved fast.
The first mover was Judge Brantley Starr in the Northern District of Texas. In the summer of 2023, weeks after Mata, he issued a standing order titled Mandatory Certification Regarding Generative Artificial Intelligence. It required every attorney appearing before him to file a certificate attesting either that no portion of any filing was drafted by generative AI, or that any language drafted by generative AI was checked for accuracy by a human being using print reporters or traditional legal databases. He explained his reasoning in plain terms. These platforms, he wrote, are prone to hallucinations and bias, and they do not swear an oath to uphold anything.
Others followed with variations. Judge Michael Baylson in the Eastern District of Pennsylvania issued a standing order requiring disclosure of AI use in any filing and certification that a human verified the accuracy of AI-generated text, including citations. Magistrate and district judges across the country adopted their own versions. Some require disclosure. Some require certification. Some ban unsupervised AI use in drafting outright. The Fifth Circuit floated a proposed rule requiring attorneys and pro se litigants to certify that AI-generated material was reviewed for accuracy, then, after public comment, decided existing rules already covered the ground and declined to adopt a special AI rule. That decision, in 2024, told you something. The court concluded that Rule 11 and the duty of candor already required exactly what the panic wanted a new rule to require.
That is the point I keep coming back to. The standing orders are useful as reminders. They put lawyers on notice in writing, which makes the eventual sanction harder to dodge. But none of them created a new duty. The duty to verify your authority is as old as the profession. What the orders did was strip away the excuse. After you have signed a certification promising a human checked the cites, you cannot stand up and say you did not know you were supposed to.
By 2025 the terrain had layers. Federal judges with individual standing orders. State courts issuing guidance. Some state supreme courts and bar bodies studying the question. Court systems in various states published advisories on generative AI in filings. The through-line in nearly all of them was the same command dressed in different clothes. If you put it in front of a court, you own it. A machine cannot own it for you.
this was always a candor problem, not a tech problem
Let me be blunt about the framing, because the tech press keeps getting it wrong.
Fabricated citations are not an AI problem. They are a candor problem and a competence problem, and both of those predate the transistor. Lawyers cited cases they never read long before ChatGPT existed. They cribbed from a treatise, trusted a summer associate, copied a string cite out of an old brief without pulling the cases. Sometimes the cases had been overruled. Sometimes they stood for the opposite of the proposition. The sin is ancient. AI just industrialized it and made it faster and more convincing.
The governing rules were already sitting right there. Model Rule 3.3 is candor toward the tribunal. It says a lawyer shall not knowingly make a false statement of law to a court, and it imposes a duty to correct false statements the lawyer has made. Model Rule 1.1 is competence, which the ABA has interpreted since 2012 to include technological competence. Rule 11 of the Federal Rules of Civil Procedure requires that legal contentions be warranted by existing law and that the lawyer conduct an inquiry reasonable under the circumstances. Feeding a prompt into a chatbot and copying the output is not a reasonable inquiry. It is the opposite of inquiry. It is delegation to a system that has no idea what truth is.
In July 2024 the ABA Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512 on generative AI tools. It walked through the duties. Competence requires understanding the benefits and risks of the specific tool, including its tendency to generate plausible but false output. Candor requires verifying anything the tool produces before it goes to a court. Confidentiality requires care about what client information gets typed into a third-party system that may train on it. Supervision requires that partners establish policies and that everyone, including nonlawyers, gets checked. Fees require that you not bill a client for hours the machine saved you as if you had done the work yourself. None of that was new law. It was old duties applied to a new gadget.
Here is where my courtroom instinct kicks in. I do not trust anything that has not survived cross-examination, and a chatbot cannot be cross-examined. It cannot be put under oath. It has no memory of its sources it can be impeached with. It generates the next plausible token based on statistical patterns, and it does that whether or not the underlying proposition is true, because it does not model truth. It models plausibility. Plausibility is exactly what a good liar produces on the stand. My entire career has been about the gap between plausible and true. That gap is the whole job.
When a lawyer files a hallucinated case, he has done the thing I spend depositions trying to expose. He has presented plausibility as truth without testing it. The court is entitled to assume he tested it. That assumption is the load-bearing wall of the whole adversarial system. Every brief carries an implied representation that a human lawyer read the authority and believes it says what the brief claims. Break that, and you are not just embarrassing yourself. You are corroding the thing that lets judges rule on paper without personally reshepardizing every cite in every filing.
Judge Castel understood this. His Mata opinion was not about robots. It was about the trust that keeps the machinery of justice moving. He wrote that the lawyers abandoned their responsibilities. That is the sentence that matters. Not the word ChatGPT. The word responsibilities.
the guardrails that survive cross-examination
So what actually works? I have talked to litigation partners, general counsel, and a couple of AI vendors trying to sell to law firms, and I have watched what real shops adopted between 2023 and 2025. Some of it is theater. Some of it is real. Here is how I sort it.
The single guardrail that matters most is the oldest one. Pull every case. Read it. Not the AI summary of it, not the headnote, the opinion. If a citation appears in a document going to a court, a licensed human has laid eyes on the actual decision and confirmed it exists, it says what the brief claims, and it has not been reversed or overruled. That is not a new best practice. It is the practice. Firms that took Mata seriously simply reinstated a discipline that had gotten sloppy over a generation of copy-paste string cites.
Second, the smart firms treated generative AI as a starting point that produces a draft, never an authority that produces a citation. There is a real difference between using a large language model to rough out an argument structure or to summarize a document you already possess, and using it to find law. The first is a drafting aid. The second is malpractice waiting to happen, because a general-purpose chatbot does not have reliable access to a verified case database and will happily invent one. Some firms banned outright the use of consumer chatbots for legal research and permitted only tools connected to real, verified databases with citation checking built in.
Third, and this is where I get skeptical, the vendors started selling AI legal-research products that promise to be grounded in real case law and to reduce hallucinations. Better than a naked chatbot, yes. Hallucination-free, no. A Stanford study in 2024 tested several major legal-AI research tools and found that even the ones marketed as reliable produced meaningful error rates, including fabricated or misdescribed authority, in a significant share of queries. The vendors disputed the methodology. Fine. But the lesson holds. A tool that hallucinates one time in six is not a tool you file from without checking. Reduced risk is not zero risk, and a court does not grade on a curve.
Fourth, firms wrote actual policies and made people sign them. Not a poster in the break room. A written AI-use policy specifying which tools are approved, what may never be entered into them, who is responsible for verification, and what the consequences are for filing unverified output. Morgan & Morgan’s post-Wyoming memo was reactive, but plenty of firms got there proactively. The good policies put a named human on the hook for every filing and made clear that the phrase the computer did it is not a defense, it is a confession.
Fifth, supervision got teeth. Under Rule 5.1 and 5.3, partners are responsible for the conduct of associates and nonlawyer staff. If a first-year runs a research query through an unapproved tool and a partner signs the brief, the partner owns it. The firms that adapted built verification into the review workflow, so that no cite reaches a signature block without a second set of eyes confirming it against a primary source. Redundancy is not inefficiency here. It is the point.
Sixth, and I like this one, some litigators started running opposing counsel’s cites through the same skeptical process, because if the other side filed hallucinated authority, that is now a weapon. I have seen briefs that flagged a nonexistent case cited by the opponent and asked the court to draw the obvious conclusion about the reliability of everything else in the filing. That is good lawyering. The Mata-style failure by your adversary is a gift. Take it.
What does not work is a certification with nobody behind it. A signed AI-disclosure form is worthless if the lawyer signing it did not actually pull the cases. The standing orders can require the certificate, but the certificate is only as honest as the person filing it. That is why I keep saying the technology is a sideshow. The certificate does not verify anything. A human verifies. The certificate just records that a human claims to have done so, and makes the sanction cleaner when the claim turns out to be false.
where I come down
I am not a Luddite about this. Used with discipline, these tools can draft a serviceable first pass faster than a tired associate at midnight. They can summarize a deposition, organize an exhibit list, surface an argument you had not framed. I have watched good lawyers use them well. The productivity is real.
But I have also spent a career learning that the fastest way to lose a case is to present something you did not personally test as if it were proven. That is what every sanctioned lawyer in this growing docket did. They presented untested output as verified authority. The tool changed. The failure did not. It is the same failure as vouching for a witness you never prepped, or reading a document into evidence you never read yourself.
The lesson of Mata v. Avianca, of Park v. Kim, of the Cohen fiasco, of Wadsworth, of Coomer v. Lindell, of Kohls v. Ellison, and of the two-hundred-plus entries piling up in that database, is not that AI is dangerous. It is that lawyers who outsource their judgment get caught, and the catching is public and permanent. The judges figured this out immediately. Castel figured it out in June 2023. He did not need a new rule. He needed lawyers to remember what their signature on a brief has always meant.
Here is the standard I use, and it has not failed me in two hundred trials. If I cannot stand in front of the judge, look her in the eye, and swear on my license that I personally confirmed every case I cited exists and says what I claim, then the brief does not go out. No machine gets to make that representation for me. No certificate substitutes for it. The day I let a statistical text generator vouch for my authority is the day I have stopped being a lawyer and started being a stenographer for a very confident liar.
Pull the case. Read the opinion. Sign your name only to what you know. Everything else is just waiting for your name to show up in the next sanctions order.
