The Third Circuit’s ROSS Decision Is Not a General AI-Training Rule

Editorial legal-tech workspace with an appellate opinion, legal research books, training-data notes, and a laptop showing abstract search results.

The Third Circuit’s ROSS Decision Is Not a General AI-Training Rule

The Third Circuit has issued one of the most important appellate opinions yet involving copyrighted material and AI training. It is also an opinion that should not be summarized too broadly.

In Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., the court held that Thomson Reuters’s Westlaw headnotes were original enough to receive copyright protection and that ROSS’s use of those headnotes to train a competing legal-research platform was not fair use. On a certified interlocutory appeal, the court affirmed the district court’s order granting Thomson Reuters partial summary judgment.

But the case involved a specific product, a specific kind of copied material, and a specific competitive use. ROSS’s system was not generative. It was designed to return passages from existing judicial opinions in response to legal questions. The court expressly described the dispute as an ordinary copyright case and distinguished questions raised by generative-AI systems that produce original expression.

That limitation is the decision’s central point for companies building or licensing AI systems: the copyright analysis does not turn on the word “training” alone. It turns on what was copied, why it was copied, what the resulting product does, and what markets the use affects.

The short answer

  • The Third Circuit affirmed partial summary judgment for Thomson Reuters.
  • The court held that 2,243 Westlaw headnotes at issue were independently created and contained enough “creative spark” to qualify as original works under copyright law.
  • ROSS’s contractors used thousands of Westlaw headnotes to help create approximately 25,000 training memoranda for a non-generative legal-research platform.
  • The court treated ROSS’s use as highly commercial and, at most, minimally transformative because both Westlaw and ROSS used the material to help users find responsive legal research.
  • The court held that ROSS’s copying was not fair use even though the underlying judicial opinions were freely available.
  • The opinion does not establish a categorical rule that training a generative-AI model on copyrighted material is unlawful or never fair use.

What ROSS copied and built

Westlaw’s headnotes are editorial annotations that identify and explain points of law in judicial opinions. Thomson Reuters’s editors choose which legal points to include and draft each headnote to be concise, accurate, and understandable on its own. Thomson Reuters does not claim copyright in the underlying opinions prepared by government officials.

ROSS was building a competing legal-research product. Its system would respond to plain-language legal questions with relevant passages from judicial opinions. The platform did not generate new expression; it returned preexisting passages.

To train the system, ROSS enlisted LegalEase Solutions, whose subcontractor Morae Global also supplied memo drafters, to prepare approximately 25,000 legal memoranda. Each memo paired a legal question with four to six judicial-opinion passages labeled “great,” “good,” “topical,” or “irrelevant” based on responsiveness. The writers used Westlaw headnotes because they provided an easy way to frame the questions. The resulting material was converted into machine-readable form and used to train ROSS’s system to identify which judicial passages responded to a user’s question.

The district court concluded that, for 2,243 headnotes, no reasonable juror could find that the corresponding memo questions had not been copied from the headnotes. On certified interlocutory review, the Third Circuit held that all 2,243 were independently created and sufficiently original, and it evaluated fair use on the undisputed facts as a matter of law.

Headnotes are not the law—but they can still be protected

ROSS argued that protecting the headnotes would give Thomson Reuters a monopoly over the law. The Third Circuit rejected that framing.

The court emphasized that the headnotes are not law. The judicial opinions remain available for anyone to read, publish, analyze, or use. But a private publisher’s editorial expression about those opinions can be protected when it reflects independent choices and a modest amount of creativity.

The court also rejected a merger-doctrine argument. Because points of law can be expressed in many ways, protecting the headnotes’ particular expression does not give Thomson Reuters control over the underlying judicial opinions or the legal ideas they convey.

The result is a distinction companies should preserve in their data reviews: public or uncopyrightable source material does not automatically make every privately prepared annotation, summary, classification, or editorial transformation free to copy.

Why the fair-use analysis went against ROSS

The court evaluated the four statutory fair-use factors together. The second factor—the nature of the copyrighted work—slightly favored fair use because the headnotes were published and more factual than fictional. The other three factors weighed against ROSS.

Purpose and character

ROSS’s use was commercial. It intended to sell a legal-research platform at prices comparable to Westlaw and to compete for the same customers.

The court also found little transformation in the relevant sense. Thomson Reuters used the headnotes to help researchers find and understand judicial opinions containing relevant points of law. ROSS used the headnotes to train a system that would help users find responsive passages in judicial opinions. The intermediate step of training an AI system did not change the ultimate purpose enough to carry the first factor.

The court contrasted ROSS’s use with software cases in which copying was necessary to access unprotected functional elements or make software operate with existing systems. ROSS already had access to the underlying judicial opinions. It copied the headnotes because they were an easier way to formulate training questions. The court’s line was direct: ease is not necessity, and ease is not a justification for copying.

Amount and substantiality

For the adjudicated 2,243-headnote subset, each headnote was treated as an individual copyrightable work; as the court put it, “for each headnote taken, ROSS copied an entire work.” ROSS nevertheless argued that the material taken represented only 0.08% of Thomson Reuters’s roughly 28 million headnotes.

The court found that percentage insufficient. Because the use was, at most, minimally transformative and ROSS could train from the freely available opinions, the amount taken was not reasonable in relation to ROSS’s purpose.

Market effect

The fourth factor also weighed against ROSS.

The court considered the legal-research platform market, where ROSS sought to compete directly with Westlaw. It also considered the developing market for licensing headnotes as AI-training data. Thomson Reuters was already using its headnotes in an AI-powered search product, and the court found that unauthorized copying could take away the opportunity to license those materials for training.

The court did not require Thomson Reuters to show an established standalone market for headnotes. Copyright law can account for harm to the value of a work and to potential derivative markets, including markets that are developing rather than fully mature.

ROSS argued that its product would increase public access to the law and that the decision could hinder AI development. The court noted that the underlying opinions were freely available, ROSS’s product was priced comparably to Westlaw, and ROSS offered no evidence supporting the broader claim that the decision would halt AI development. AI’s presence in a product did not give ROSS carte blanche to copy copyrighted material.

The generative-AI boundary matters

The opinion includes a footnote addressing arguments made by the Department of Justice in separate generative-AI litigation. The DOJ argued that training a large language model capable of generating original responses can be transformative. It separately argued that the training there did not result in substitutive competition.

The Third Circuit said those concerns did not apply to ROSS. ROSS’s platform could not generate original expression, and the evidence showed that it was trained to create a commercial substitute for Westlaw. Accordingly, the opinion does not decide whether training a generative model capable of original expression is transformative or how an absence of substitutive competition bears on fair use.

That is not a throwaway qualification. It prevents the opinion from being used as a simple appellate declaration that all AI training on copyrighted works fails fair use. It also prevents companies from treating the decision as irrelevant merely because their system is generative. The opinion identifies the factual and market questions that will likely matter in later cases.

What companies should take from the decision

Companies developing or licensing AI systems should review at least five questions:

  1. What exactly entered the dataset? Separate public-domain or otherwise unprotected source material from private annotations, summaries, labels, classifications, editorial selections, and other expressive inputs.
  2. Was copying necessary or merely convenient? Document why a particular source was used and whether the same function could have been achieved from available underlying material or through a license.
  3. What does the resulting product do? A training step does not answer the fair-use question. Assess whether the product generates new expression, retrieves or reproduces source material, supports a different function, or competes in the same market as the source.
  4. What market does the use affect? Consider existing product markets and potential derivative markets that creators of original works would generally develop or license others to develop, including developing markets for training data.
  5. What evidence supports the analysis? Preserve dataset provenance, permissions, source classifications, model-development records, product comparisons, and the reasoning behind inclusion decisions.

Those controls are not a substitute for a case-specific legal opinion. They do, however, make the relevant facts visible before a dispute turns into emergency discovery.

The bottom line

The Third Circuit’s ROSS decision is a significant appellate ruling on copyright and AI training, but it is not a general rule that AI training is unlawful.

The court held that ROSS’s use of protected editorial headnotes to train this competing, non-generative legal-research product was not fair use. The 2,243 headnotes were original enough to protect; ROSS’s use shared Westlaw’s ultimate purpose; each copied headnote was treated as an entire work; and the use harmed both the legal-research market and a developing potential derivative market for licensing headnotes as AI-training data.

For AI developers and data licensors, the practical lesson is narrower and more useful than a sweeping headline: classify the material, document the purpose, test the competitive effect, and do not assume that calling a use “training” resolves the copyright analysis.

Sources

This article is general information, not legal advice.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *