PPenwell Law

Practice · Generative AI, training data and IP

Generative AI counsel: training data, model outputs and intellectual property

Generative AI companies face three recurring legal questions: what the model was trained on and under what rights, who owns and may use its output, and what happens when an output resembles someone's work or voice. Penwell Law advises on all three and builds the legal frameworks for training data, model use and model governance for companies building models, companies building on top of them, and the rights holders on the other side of the table.

Start the conversation

Training data rights and licensing

Investors and acquirers now diligence training data the way they diligence code ownership. Penwell Law documents the provenance of a training corpus, licenses datasets and catalogs where a license is available, and gives a written fair-use analysis under the Copyright Act where it is not, stating which uses current case law supports and which remain open. Penwell Law writes that analysis for review by a buyer's counsel. California's AB 2013 requires generative AI developers to publish a summary of their training data, and the AI Transparency Act (SB 942, as amended by AB 853) adds provenance and disclosure duties for AI-generated content, so the provenance record supports both disclosures.

Web scraping and the chain of custody for data

Scraped web data is the most common source of training data and the most contested. Penwell Law treats a scraping program as five questions, each with its own body of law, and answers them source by source before the first request is sent.

Access. Liability under the federal Computer Fraud and Abuse Act turns on whether the scraper passed an authorization gate, not on its purpose. California Penal Code section 502 reaches further: it requires only knowing access followed by taking, copying or using data without permission. For a public page, an IP block or a cease-and-desist letter does not create an authorization gate under the Computer Fraud and Abuse Act; behind a login, a cease-and-desist letter revokes permission and further access is unauthorized. A robots.txt file is not an access control, and Penwell Law also treats it as evidence of what the site owner wanted. Terms. A site's terms of use bind a visitor who had conspicuous notice of them and took an action showing assent, and whether an automated client that never renders a page can assent is an open question; Penwell Law drafts the customer representations and the indemnity so that the risk sits with the party that chose the target. Conduct. An accurate user-agent, compliance with rate limits and the sourcing of any proxy pool affect whether a crawl is treated as legitimate or as evasion under California's Unfair Competition Law and the fraud provision of the Computer Fraud and Abuse Act; trespass to chattels requires actual or threatened damage to the target's systems, which a rate-limited crawl does not cause. Copyright. Copyright protects authored expression, not facts, so only the authored portions of a page are protected. Fair-use decisions on training have turned on the purpose of the use, how the copies were acquired and proof of market harm, and removing copyright management information from a page creates exposure under section 1202(b) of the Digital Millennium Copyright Act when the removal furthers an infringement. Courts assess training and output separately, so a fair-use finding on training does not protect an infringing output. Personal data. Pages containing personal information bring the CCPA, CPRA and GDPR into the program, including whether the scraper acts as a controller or a processor, and the EU database right protects databases built on substantial investment, which U.S. law leaves unprotected.

Penwell Law then builds the chain of custody that diligence requires: a provenance record for each dataset showing its source, the terms or license it was collected under, the collection date, the filters applied, the handling of personal information and who approved it, so that a buyer's counsel, a regulator or a court can trace any training example to its origin. A company should not warrant that scraped pages are free of third-party rights; its indemnity should cover how the content was obtained, not a customer's later use of it. The provenance record also lets the company answer removal requests from rights holders and data subjects.

Model outputs: ownership, warranties and the reliance problem

The U.S. Copyright Office does not register output that lacks sufficient human authorship, and customer contracts must account for that position. Penwell Law drafts output terms a company can defend: what the customer owns, what the company disclaims, the limits on reliance on a model's output, and the allocation of indemnities when a third party claims an output infringes. For consumer products, the work includes age-gating, non-affiliation and output-reliance disclosures.

Likeness, voice and the DMCA

Digital replicas and synthetic voices raise right-of-publicity claims under California Civil Code section 3344 and its common-law counterpart, alongside trademark and false-endorsement exposure under the Lanham Act. Penwell Law counsels on consent and licensing frameworks for replica products, and it defends against those claims. On the platform side, it builds the DMCA section 512 notice-and-takedown process and the repeat-infringer policy that keep the safe harbor available, and the content-provenance and labeling practices that regulators and platforms expect.

Content and catalog licensing

Penwell Law also represents rights holders in these deals. It advises artists, labels, publishers and creative businesses on recording and publishing agreements, synchronization and catalog licensing, performance contracts, and royalty and rights management, and it advises the AI companies that license those catalogs. Penwell Law handles the AI versions of these questions in the same practice: licensing sound recordings and other works for training, and ownership of AI-assisted output.

Two decades of owning and licensing intellectual property

Liam J. Penwell spent close to two decades owning and licensing intellectual property before law, as a founder and operator whose production company served Google, Uber, LinkedIn, Salesforce and Twitter, and he has released and performed original work under his own names. He has taught copyright, trademark, contracts and AI-and-copyright issues at the San Francisco Conservatory of Music.

← Back to the main page

Frequently asked questions.

Is training an AI model on copyrighted material fair use?

It depends on the facts, and the case law is still developing. Recent decisions have turned on the purpose of the training, whether the copies were lawfully acquired, and market effect. Penwell Law gives companies a written analysis of their specific corpus and use under the Copyright Act's four factors, licenses data where a license is available, and documents the reasoning where it is not.

Is it legal to scrape websites for AI training data?

It depends on what is scraped, how, and from where. A site's terms of use are a contract with a visitor who assented to them; bypassing access controls can reach the Computer Fraud and Abuse Act and California Penal Code section 502; the content may be copyrighted; and pages with personal information bring the CCPA, CPRA and GDPR into play. Penwell Law reviews a scraping program source by source before it runs and builds the provenance record that shows what was collected and under what terms.

Does removing boilerplate or attribution from scraped pages create legal risk?

It can. Copyright management information such as author names, titles and copyright notices is protected by section 1202 of the Digital Millennium Copyright Act, and removing it creates liability when the removal furthers an infringement. Penwell Law reviews what a pipeline strips from a page and what it keeps, and documents the reasoning.

Who owns the output of a generative AI model?

Under current U.S. Copyright Office guidance, output without sufficient human authorship is not protected by copyright, so ownership is largely a matter of contract. The practice drafts the output terms between a company and its customers, and between the company and the model provider it builds on.

What legal issues do synthetic voice and digital replica products raise?

These products create right-of-publicity exposure under California Civil Code section 3344 and the common law, Lanham Act false-endorsement claims, and consent and licensing questions for the person whose voice or likeness is used. Penwell Law builds the consent frameworks and defends the claims.

Does a generative AI platform need a DMCA process?

A platform hosting user content needs a compliant section 512 notice-and-takedown process, a registered agent, and an enforced repeat-infringer policy to keep the safe harbor.

Liam writes on the legal questions behind building and shipping technology: AI governance, privacy, and content rights. Read the blog →

Latest: Your Crawler Needs an Alibi

New posts by email. Unsubscribe anytime.

Start the conversation.

Contact

Based in Burlingame, California, working with clients and teams across Silicon Valley, the Bay Area, and nationally.

Email
Telephone
Office Location
Burlingame, CA
Mailing Address

Send a Message