What changed
Three AI copyright dockets surfaced by the litigation watch show the operational burden behind the headline disputes.
In Concord Music Group, Inc. v. Anthropic PBC, the docket includes declarations supporting a preliminary injunction motion. The listed attachments include redacted Anthropic emails and “Anthropic Hugging Face Prompts.”
In Getty Images (US), Inc. v. Stability AI, Ltd., Getty filed a complaint against Stability AI entities with a jury demand and exhibits.
In In re Google Generative AI Copyright Litigation, the docket includes amended complaints, redlines, and exhibits, including copyright registration material and a representative list of websites alleged to contain misappropriated content.
The merits remain contested. But the signal for counsel is already concrete: these disputes are being framed through records, exhibits, declarations, prompts, internal communications, and amended pleading detail.
The hinge: evidence-readiness, not just fair-use posture
AI copyright risk is often discussed at the level of legal defenses, policy statements, or product principles. These dockets show why that is not enough.
When litigation pressure arrives, the record may need to answer narrower questions:
- What data sources were used or excluded?
- Which datasets were licensed, public, user-provided, scraped, or otherwise obtained?
- Who approved those sources, and on what basis?
- What did internal communications say about acquisition, training, model behavior, or output concerns?
- What prompt and output testing was conducted?
- What vendor or customer representations were made, and can the company substantiate them?
Those questions are operational before they are legal. If the relevant records are scattered, overwritten, informal, or inconsistent with contract language, counsel’s options narrow quickly.
Preservation steps to run now
For AI builders and enterprise users of AI systems, the immediate control is not to predict outcomes in these cases. It is to make sure the company can reconstruct its own facts.
A practical evidence-readiness review should include:
- Dataset source logs
Confirm that source logs identify where training, tuning, evaluation, or retrieval corpora came from. The log should be specific enough for counsel to distinguish categories of material rather than relying on broad labels.
- Data category mapping
Map data into operational categories such as licensed, public, user-provided, scraped, excluded, or removed. The point is not cosmetic taxonomy; it is being able to explain what was used and why.
- Licensing and exclusion decisions
Preserve the records showing why particular corpora were licensed, not licensed, blocked, filtered, or retained. If a decision was made by procurement, product, engineering, or legal, keep the decision trail connected.
- Prompt and output testing records
Retain prompt/output test sets, results, escalation notes, and remediation decisions where a system was tested for reproduction, close similarity, or other content-risk concerns. Concord’s docket reference to Hugging Face prompts is a reminder that prompts themselves can become evidence.
- Internal communications hold points
Identify channels where data acquisition, model behavior, or output issues are discussed. Preservation should cover informal channels if those are where relevant decisions were actually made.
- Version-aware product records
Keep enough version history to connect a complaint, customer issue, or test output to the model, dataset, or configuration in use at the relevant time.
Contract and procurement questions
For enterprise buyers, the same signal applies through vendor diligence and contract files. A vendor representation is only useful if it maps to a record the vendor can support.
Procurement and product counsel should pressure-test:
- What does the vendor say about training-data sources?
- Does the vendor distinguish licensed, public, user-provided, scraped, and excluded data?
- What records does the vendor keep about prompt/output testing?
- Are contractual representations aligned with what the vendor can actually prove?
- If the vendor receives a copyright claim, what information will the customer receive about affected models, datasets, or outputs?
- Are customer-facing product claims broader than the underlying diligence file supports?
For AI builders, the same questions should be asked in reverse. Do sales, security, procurement, and legal teams use the same description of data provenance and model behavior? If not, litigation may expose the inconsistency.
Escalation trigger
Move from general policy review to evidence-readiness review when either of these conditions is present:
- the product uses third-party corpora in training, tuning, evaluation, retrieval, or output controls; or
- testing shows outputs that could be alleged to reflect protected material.
At that point, counsel should not rely only on statements of principle. The file should contain the underlying records: data-source documentation, licensing decisions, prompt/output tests, internal escalation notes, and vendor or customer representations.
Caveats
These are docket signals, not final rulings. Complaints, amended complaints, declarations, exhibits, and preliminary-injunction materials contain allegations and advocacy positions; they do not establish liability or predict how the courts will resolve the merits.
The sources are also litigation-watch docket materials and excerpts. They are useful for identifying the types of records becoming focal points, but they are not a complete record of each case.