PPenwell Law

← Blog

Your Crawler Needs an Alibi

Zero data retention is the right promise for an AI data company, and enterprise buyers want it. The promise serves both sides until a regulator or site operator decides it has a problem. The good news? A vendor can keep that promise and still show what its crawler did, with an alibi so airtight even Benoit Blanc would be impressed.

Read literally, zero retention leaves the vendor with nothing to produce when a site operator or a regulator asks how a page was acquired. The obligation reaches the page content the crawler retrieves and the output the model returns. It does not reach a record of the fetch itself: the exclusion rule the crawler matched, the identity the crawler presented, and a signature over those fields. A vendor that keeps that record and discards the content satisfies the buyer's addendum and can still document, for each page, that its crawler honored the operator's rules.

The question below is a composite of the questions founders of AI data companies bring to counsel.

Our enterprise deals keep stalling in security review. We have the SOC 2 report, the data processing addendum, the zero-retention addendum and a no-training clause. Procurement still will not sign. What are we missing, and what needs to change?

The contracts in that stack allocate a loss once the facts are properly presented. If the buyer’s privacy reviewer cannot find the retention and training terms in the addendum, or worse, litigation counsel cannot see what backs the acquisition indemnity, well, what did you expect?

The SOC 2 report attests to the vendor’s controls. Fine. It’s the zero-retention addendum and the no-training clause that sit outside the data processing addendum, which is the document the buyer’s privacy reviewer actually builds an Article 28 file around. None of the four documents records whether the vendor honored a site's exclusion rules on a given day or what identity its crawler presented, and those are exactly the facts a vendor has to unravel, and fast, when a site operator disputes a fetch and the customer calls on the indemnity. Good luck parsing that one without a record.

So, what do we do? First, move the zero-retention and no-training promises into the data processing addendum and a retention annex. Next, keep a signed provenance record of every fetch that holds no page content and no output. With that record the vendor keeps its zero-retention promise and can still prove how it acquired data. Feeling groovy yet?

Third, where the vendor selects the sources, use those records to support the section 3111 disclosures of customers that develop or substantially modify generative AI systems made publicly available to Californians. Civil Code section 3111 requires those developers to post a high-level summary of their training datasets, including the sources or owners of the datasets, and a report drawn from signed records documents that part of the summary. My earlier post on provenance covered the audit trail an AI system leaves behind without planning; this post covers the record a vendor builds on purpose.

The indemnity turns on a fact neither party can prove

Take the most careful allocation counsel can draft for a company whose product fetches, retrieves or ingests third-party data. The vendor warrants that it honored each site's exclusion rules, presented an honest identity and stayed within its allotted request rate. The customer warrants that it had a basis to target each source and will not redistribute the results. Directional indemnities back each warranty, with a concurrent-causation clause for the overlap.

Then a site operator sues and asks two questions: did the vendor honor its exclusion file on March 14, and what identity did the crawler present? The vendor and the operator give opposite answers, and the indemnity apportions the loss according to whichever answer is defensible. The operator's server logs show requests arriving, not which exclusion rule the crawler applied, and the vendor holds nothing signed to answer them. I have watched smart engineers try to reconstruct a March 14 from memory. It does not go well. Most of the fields needed to answer the question already pass through the vendor's pipeline; they just need to be leveraged properly.

Between vendor and customer, the record makes the indemnity apportionment work. The record belongs in the indemnity and in a records schedule to the master agreement, where the customer can bargain over its terms. Against the site operator, which signed nothing, the record is the vendor’s evidence. A contract cannot bind the operator, but the master agreement can bind the vendor. That is where the vendor warrants that it will generate the record, certify its accuracy, and retain and produce it on demand, no matter which site operator decides to make a federal case out of a robots.txt file.

What a data processing addendum covers

Article 28(3) of the GDPR sets the terms on which a processor handles personal data for a controller. Article 4(2) includes collection in the definition of processing, so a vendor whose crawler fetches pages containing personal data on a customer's instructions is a “processor” within Article 28 from the moment of the fetch.

Under those terms the processor acts only on documented instructions, keeps the data confidential, secures it, flows its obligations down to sub-processors, assists with data-subject requests, deletes or returns the data at the end of the engagement, and makes available "all information necessary to demonstrate compliance." California reaches the same relationship by definition: Civil Code section 1798.140(ag) defines a service provider as one bound by a contract that prohibits selling or sharing the personal information; retaining, using or disclosing it outside the business purposes specified in the contract; retaining, using or disclosing it outside the direct business relationship between the service provider and the business; and combining it with personal information from other sources, except as the regulations permit.

Neither regime requires a record of how the vendor acquired a page from a third-party site. The omission costs little in conventional SaaS, where the personal data belongs to the customer's own users and the facts are seldom contested. A web-data vendor's exposure arises at the fetch from the third-party site, which is the step the standard addendum does not address.

Where the training and retention promises sit

When a buyer’s general counsel says the company cannot send its data to an AI vendor, the worry is that the vendor will train on the data, keep the data, or let the data reach another customer. These are separate issues and promises, which, in typical fashion, live in different documents.

I reviewed the public documents of two widely deployed model providers in September 2026. Anthropic’s Commercial Terms of Service, effective June 17, 2025, state in Section B that “Anthropic may not train models on Customer Content from Services,” and Section C incorporates its Data Processing Addendum by reference. OpenAI’s Services Agreement, effective January 1, 2026, states in Section 4.2 that “OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.” Both commitments sit in the providers’ commercial terms. Neither provider’s data processing addendum states a no-training commitment.

Both addenda address deletion at the end of the relationship. Anthropic commits in its Data Processing Addendum, effective February 24, 2025, to delete customer data within thirty days of termination, subject to stated exceptions. In Schedule 1 of its Data Processing Addendum, effective January 1, 2026, OpenAI sets the duration of processing as “the Term and such time required thereafter for the Parties to perform their applicable obligations following the end of the Term, including data deletion.” OpenAI mentions training in that addendum only once, for staff who handle sensitive data. Neither addendum states a zero-retention term. OpenAI describes its Zero Data Retention endpoints as offered “under our ZDR amendment.”

A buyer builds its Article 28 compliance file around the data processing addendum, because that is where the terms Article 28(3) requires ordinarily sit. A no-training covenant in the master agreement binds the vendor. A vendor that repeats the covenant in the addendum’s provision on documented instructions puts the promise in the right place.

Where a vendor states a retention period at all, the period ordinarily appears in product documentation or a configuration setting. OpenAI drafted its addendum to make that setting the customer’s decision. Section 2.1 treats instructions given through the service’s configuration tools as the customer’s documented instructions. Section 3.3 makes the customer responsible for configurations such as retention periods and deletion.

The European Data Protection Board lists “for how long shall they be processed” among the essential means of processing reserved to the controller, in Guidelines 07/2020 at paragraph 40. The Board treats a customer that chooses a standardized, take-it-or-leave-it service as the controller, because the customer decided to use that service for its own purposes. A vendor that offers zero data retention should write it into a retention annex, or into a setting the addendum names as a documented instruction, so the contract records the retention period as the customer’s decision.

Zero data retention and the evidence a vendor still needs

A vendor that keeps no customer content after a request completes has no stored content to reach another tenant or to produce under compulsion. Enterprise buyers demand zero data retention for that reason. The facts of how the vendor acquired a page are separate from the page itself, so a zero-retention vendor that needs to prove it honored a site’s rules on March 14 needs a separate record of the fetch.

In the consolidated copyright litigation against OpenAI in the Southern District of New York, In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143, the magistrate judge ordered OpenAI in May 2025 to preserve and segregate output-log data that would otherwise have been deleted. OpenAI stated that the order did not affect API customers using its Zero Data Retention endpoints, because that content is never stored. Logs that OpenAI had retained in the ordinary course of business were discoverable. In January 2026 the district judge affirmed orders compelling OpenAI to produce a sample of twenty million of those logs, de-identified. The orders arose in one case and concern output logs, which differ from acquisition records.

A vendor that holds a signed acquisition record can put the facts of each fetch before a court without retaining the page content or output its zero-retention promise covers. Nice.

The provenance record

A provenance record keeps the facts of each fetch and none of the page content or output, exactly how it should be. For every retrieval the vendor writes down what it fetched, when, from which network and under what identity, which exclusion rule applied, and what came back, reduced to a hash. The retention annex should disclose that the target URL, the only customer input in the record, is kept. Query strings can carry search terms, session tokens or personal data, so watch out for those. The vendor should strip or hash any parameter its exclusion decision does not need, consistent with the data-minimization principle in Article 5(1)(c). The buyer keeps the zero-retention commitment it negotiated and has an audit trail for its litigation counsel.

Most of these fields come from systems the vendor already runs. The target URL and the timestamps fix the operation. The HTTP status and byte count separate a successful fetch from a target error and settle whether a 4xx response is billable. The robots.txt evaluator knows whether it checked an exclusion file, which rule matched and what it decided. The HTTP client knows the user-agent string it sent. The network stack knows the egress IP range, the autonomous system and any proxy network the request passed through. Where the vendor uses Web Bot Auth, the signing key identifies the request. If a site answers with a challenge, a rate-limit response or a block, the record notes that response and the crawler’s next step. The account system holds any override in force and the signed instrument behind the override. The vendor keeps a SHA-256 hash of the response in place of the response itself, and signs the record with a timestamp from an append-only log. With illustrative values, a single fetch produces a record like this:

target_url:              https://example.com/team/jane-doe
requested_at:            2026-03-14T17:02:11Z
fetched_at:              2026-03-14T17:02:12Z
http_status:             200          response_bytes: 48213
robots_evaluated:        true
robots_source_url:       https://example.com/robots.txt   (fetched 2026-03-14T17:02:11Z)
robots_rule_matched:     "User-agent: *  Allow: /team/"
robots_decision:         allow
user_agent_presented:    "ExampleFetcher/2.1 (+https://example-vendor.com/bot)"
egress_ip_range:         203.0.113.0/24   egress_asn: AS64500
proxy_network:           none
request_signature:       web-bot-auth     signature_key_id: wba-2026-01
access_control_response: none             crawler_action: fetched
rate_limit_applied_ms:   1000
override_in_force:       false        override_addendum_id: none
retention_mode:          zero_data_retention
content_sha256:          9f2c...e41a
record_signature:        ...           signing_key_id: k-2026-01   log_inclusion_proof: ...

A customer that kept the returned bytes can hash them and match the record. Because dynamic pages rarely return identical bytes twice, the hash ties the record to the response delivered at that moment. A later change to the record breaks its signature or its inclusion proof. The cost to the vendor is storage, key management, and a transparency log operated correctly for years at crawl scale. YMMV.

Drafting the warranty and the evidentiary burden

The vendor should not warrant that the record is admissible. (And neither should the vendor’s counsel!) Admissibility is a question for the tribunal. A warranty of admissibility would hand opposing counsel the most direct attack on the entire schedule. I would have the vendor warrant that each record accurately reflects, in all material respects, the facts of the retrieval as the system observed them at the time. I would also have the vendor warrant that the Services generate each record automatically at the time of the retrieval and use it in their ordinary operation to apply exclusion decisions and calculate billing.

That language tracks the business-records exception in Federal Rule of Evidence 803(6), which reaches a record made at or near the time by, or from information transmitted by, someone with knowledge, kept in the course of a regularly conducted activity, where making the record was a regular practice of that activity, unless the opponent shows that the source of information or the method or circumstances of preparation indicate a lack of trustworthiness. The Advisory Committee Note to Rule 803 explains that the Supreme Court in Palmer v. Hoffman, 318 U.S. 109 (1943), upheld the exclusion of an accident report that “was prepared for use in litigating, not railroading.” Because the Note stresses records of routine operations, the warranty ties each record to exclusion decisions and billing.

Rule 902(13) makes self-authenticating “a record generated by an electronic process or system that produces an accurate result,” on certification by a qualified person. Rule 902(14) does the same for data copied from a device, medium or file and “authenticated by a process of digital identification.” Both require the advance notice set out in Rule 902(11). The Advisory Committee Note to Rule 902 states that Rule 902(13) “is solely limited to authentication, and any attempt to satisfy a hearsay exception must be made independently.” The contract should oblige the vendor to furnish, on request, a Rule 902(11) certification of the business-record foundation and a Rule 902(13) certification of the process that generates the record.

Evidentiary burden. Where a claim turns on whether the Services honored a Target’s exclusion rules, presented a particular identity, or observed rate limits at the time of a retrieval, Vendor shall produce the Provenance Record generated under Schedule P on request. For the purposes of the apportionment in this Section only, a Provenance Record so produced is prima facie evidence of those facts, and a party disputing it bears the burden of producing evidence that the record is inaccurate. This Section does not shift the burden of persuasion. Vendor’s liability for an inaccurate Provenance Record is subject to the limitation of liability in this Agreement, and where an inaccuracy is not a breach of the Provenance Record warranty, Vendor’s sole obligation is to correct the record.

California stops its own presumption for computer records at the same point. Evidence Code section 1552 presumes that a printed representation of computer information accurately represents that information, and makes the presumption one affecting only the burden of producing evidence. A vendor can concede an adverse consequence for a missing record as a negotiated trade. The schedule should confine that consequence to records within the vendor’s control that were lost other than under a documented retention schedule, a third-party system failure or force majeure.

Controllership of the provenance record

Target URLs reach the vendor with the customer’s instructions, as data the vendor processes for the customer. A vendor that decides for itself to keep those URLs in a signed, tamper-evident dataset outside the customer’s deletion elections determines the purposes and means of that processing, which makes it a controller under Article 4(7). Article 28(10) treats a processor that does so in breach of the Regulation as a controller, so the data processing addendum should expressly recognize the vendor as an independent controller of retrieval metadata. A target URL can identify a person, and the vendor alone decides whether to create the record, what the record contains and how long to keep the record. In California, 11 CCR section 7050(a)(5) permits a service provider to retain, use and disclose personal information for the purposes the regulation incorporates from Civil Code section 1798.145(a), which include exercising or defending legal claims under section 1798.145(a)(1)(E).

The schedule names the vendor as an independent controller of the provenance records, relying on its legitimate interests in operating the service and in establishing, exercising and defending legal claims under Article 6(1)(f). Operating the service includes compiling domain-level sources reports for customers, which contain no personal data. The vendor processes the records for those purposes only, does not enrich or resell the records, applies its security annex to them, lists the activity in its Article 30 register, and discloses the activity in its public privacy notice. The schedule states the retention period per claim type as the record’s operational life for exclusion decisions and billing, extended to the limitation period for claims about the fetches it records, and reviews the period when that limitation period changes. The storage-limitation principle in Article 5(1)(e) allows identifiable data to be kept no longer than is necessary for the purposes of the processing. The vendor declines erasure requests directed at a record to the extent Article 17(3)(e) preserves processing necessary for legal claims.

People identifiable from a target URL are strangers to both parties. The European Data Protection Board’s draft Guidelines 03/2026 on web scraping for generative AI say the disproportionate-effort exemption in Article 14(5)(b) “should not be routinely relied upon” outside archiving, research and statistical processing, and call for a balancing of the controller’s effort against the impact on data subjects. The schedule should record that balancing before the vendor relies on the exemption. The vendor’s public notice should give a precise indication of the sources and describe the crawler’s characteristics, as paragraph 30 of the draft guidelines advises. The vendor answers objections under Article 21(1), which lets processing continue for the establishment, exercise or defense of legal claims. The schedule states that the arrangement is independent controllership and not joint controllership under Article 26, because the vendor alone decides to create, sign and keep the record, for purposes the customer does not pursue, even though the customer’s choice of targets determines which fetches the record describes. The Court of Justice held in Koninklijke Nederlandse Lawn Tennisbond, Case C-621/22 (October 4, 2024), that a legitimate-interest basis requires a lawful interest, processing strictly necessary for that interest, and data-subject interests that do not override it, and confirmed that a commercial interest can qualify. The schedule should therefore state the purpose narrowly.

Crawler identity and the exclusion warranty

If a crawler checks a site’s exclusion rules against one product token and then sends a randomized, browser-mimicking user-agent, a signed record of that fetch is a contemporaneous, cryptographically attested admission that the identity it presented did not match the identity it checked. Identity now extends to the network. Cloudflare de-listed Perplexity as a verified bot in August 2025 after observing requests from IP addresses outside Perplexity’s published ranges, rotated across autonomous systems, and a generic browser user-agent used when Perplexity’s declared crawler was blocked. The IETF has chartered a Web Bot Authentication working group to standardize cryptographic authentication for automated clients. The group has not yet published an RFC. Site operators now also plead circumvention of blocks under section 1201 of the Digital Millennium Copyright Act. Reddit sued SerpApi, Oxylabs, AWMProxy and Perplexity in Reddit, Inc. v. SerpApi LLC, No. 1:25-cv-08736 (S.D.N.Y.). On July 31, 2026, the court largely denied the motions to dismiss filed by SerpApi and Perplexity and sustained the section 1201(a)(1)(A) circumvention claims against both. In Google LLC v. SerpApi, LLC, No. 4:25-cv-10826 (N.D. Cal.), the court on July 20, 2026 dismissed without leave to amend Google’s claims over search results containing no copyrighted content and gave leave to amend the claims over results with a copyrighted component. Google filed an amended complaint in August. The build order is a stable, documented, blockable product token and published IP ranges first, then the warranty that exclusion rules and access controls are honored, then the record.

The web ran on robots.txt for close to three decades before the Robots Exclusion Protocol became RFC 9309 in 2022. RFC 9309 is still a Proposed Standard. A warranty of conformance should track the specification’s own normative levels, since RFC 9309 distinguishes MUST from SHOULD. RFC 9309 does not define a crawl-delay directive, so the warranty should list by name any directive outside the RFC that the vendor honors, such as crawl-delay. Cloudflare’s Content Signals Policy, released under a CC0 license in September 2025, adds search, ai-input and ai-train directives. Really Simple Licensing publishes machine-readable licensing terms supported by Reddit, Yahoo, Fastly, O’Reilly and others. The IETF’s AI Preferences working group has not published its specifications as an RFC as of September 2026.

In the European Union, Article 4(3) of the Digital Single Market Directive, Directive (EU) 2019/790, lets a rightholder reserve text-and-data-mining rights “in an appropriate manner, such as machine-readable means in the case of content made publicly available online,” and recital 18 counts “metadata and terms and conditions of a website or a service” as machine-readable means. In Kneschke v. LAION, the Hanseatic Higher Regional Court in Hamburg ruled for LAION on December 10, 2025 (5 U 104/24). The court held that the reservation of rights on the image agency’s website, when LAION downloaded the photograph, lacked the machine-readable form that section 44b(3) of the German Copyright Act requires. The court held separately that the use was permitted under the scientific-research exception in section 60d, which covers research organizations acting on non-commercial terms, cultural heritage institutions and individual non-commercial researchers. A commercial data vendor cannot rely on that ground. The Federal Court of Justice heard the appeal on September 3, 2026, issued no decision, scheduled its decision for December 17, 2026, and indicated that the case turns on questions of EU law. The Board’s draft Guidelines 03/2026 note that legitimate interest is often used by private entities for scraping, apply its three cumulative conditions, and list safeguards for collection and later processing. The guidelines are a consultation draft and bind no one. A vendor that emits a signed acquisition record already documents the collection facts they describe. A warranty to honor “applicable signals” that names no signal gives the customer nothing to enforce. The warranty should name each standard and date each version.

California’s AB 2013 training-data disclosures

The same records support the written disclosure every developer that makes a generative AI system publicly available to Californians must now post. Civil Code section 3111, added by AB 2013, requires the developer of a generative AI system released on or after January 1, 2022 and made publicly available to Californians for use to post documentation of the data used to train the system, including “[a] high-level summary of the datasets.” Developers had to post the first documentation on or before January 1, 2026, and must post new documentation before each later public release or substantial modification. The documentation covers twelve items, including the sources or owners of the datasets, whether they include copyrighted, trademarked or patented material, whether they were purchased or licensed, whether they include personal information, and the period of collection.

Section 3111 does not excuse data supplied by a third-party vendor, so a customer that fine-tunes a system and makes the modified system publicly available to Californians owes the same disclosure for vendor-sourced content as for content it collected itself.

Its only exemptions, in subdivision (b), cover systems whose sole purpose is security and integrity, the operation of aircraft in the national airspace, or national security, military or defense purposes made available only to a federal entity. The statute is in force. In X.AI LLC v. Bonta, No. 2:25-cv-12295 (C.D. Cal.), the district court denied xAI’s motion for a preliminary injunction on March 4, 2026. xAI’s appeal is pending in the Ninth Circuit as No. 26-1591.

Outside California, Article 53 of the EU AI Act requires providers of general-purpose AI models placed on the market from August 2, 2025 to publish a training-content summary, and gives providers of earlier models until August 2, 2027. The Digital Omnibus on AI, Regulation (EU) 2026/1744, left that obligation in place.

Any customer that trains on vendor-sourced data will be asked where that data came from. Where the vendor selects the sources, through search, crawl expansion or prebuilt datasets, a vendor with signed provenance records can answer with a sources report compiled from them, listing sources or owners and collection periods at the domain level, without handing over the records or keeping any page content or output. The master agreement should attach the records schedule alongside the data processing addendum, so the zero-retention terms and the acquisition record appear in the same contract.

What the record does not establish

Whether a vendor had the right to fetch a page depends on contract, access law and privacy law. The provenance record answers none of those questions. Under section 1798.140(v)(2), personal information excludes both publicly available information and lawfully obtained, truthful information that is a matter of public concern. “Publicly available” is a defined term with three limbs: information lawfully made available from federal, state or local government records; information a business has a reasonable basis to believe is lawfully made available to the general public by the consumer or from widely distributed media; and information made available by a person to whom the consumer disclosed it, if the consumer did not restrict it to a specific audience. Whether a given page falls within either exclusion is a question the provenance record does not answer, and biometric information a business collects about a consumer without the consumer’s knowledge is never publicly available.

The federal access cases leave the right to fetch unsettled as well. Van Buren v. United States, 593 U.S. 374 (2021), confined the Computer Fraud and Abuse Act to a gates-up-or-down inquiry, left open whether a contract or policy can close a gate, and says nothing about contract or copyright claims. hiQ Labs v. LinkedIn, 31 F.4th 1180 (9th Cir. 2022), is a preliminary-injunction ruling under a serious-questions standard, and the litigation ended in December 2022 when hiQ consented to judgment and a permanent injunction. On contract, the court in Meta Platforms, Inc. v. Bright Data Ltd., No. 3:23-cv-00077 (N.D. Cal. Jan. 23, 2024), granted summary judgment to Bright Data because Meta’s terms did not reach logged-off scraping of public data. How data was acquired can decide liability. In Bartz v. Anthropic PBC, No. 4:24-cv-05417 (N.D. Cal.), the court held on June 23, 2025 that training on lawfully acquired books was fair use and denied summary judgment on copies downloaded from pirate libraries. The court granted final approval of a $1.5 billion class settlement on July 20, 2026. The provenance record proves what a company did. The contract should place the representation that the company was entitled to fetch each target on the party that selected the target.

Before the next enterprise redline

First, move the zero-retention and no-training promises into the data processing addendum, and put the retention terms into a retention annex as a documented instruction.

Second, build the stable identity and published network ranges, then the warranty that exclusion rules and access controls are honored, then the signed record and its schedule. The schedule carries the evidentiary-burden clause and the independent-controller paragraph, drafted before a buyer’s counsel drafts them.

Third, where the vendor selects the sources, give customers that owe a section 3111 disclosure a sources report compiled from the records.

Frequently asked questions

Is a signed record simply more discoverable evidence against the vendor? It is discoverable, as the pipeline logs it draws on already are. A record the vendor designs carries a warranty the vendor controls. A pipeline that honors exclusion rules and presents an honest identity already generates most of the fields. The alternative is a subpoena answered from scattered logs that lack the matched rule and the override status and carry no warranty. Because a signed record also exposes a crawler that misrepresents its identity, identity comes first in the build order.

Does a provenance record conflict with zero data retention? No. The record contains no page content and no output, so the customer’s content stays deleted. It does contain the target URL the customer submitted. The retention annex should define customer content to exclude retrieval metadata and state that the vendor keeps that metadata as an independent controller for the period set in the schedule. In California, section 1798.145(a)(1)(E) provides that the statute does not restrict a business’s ability to exercise or defend legal claims. Under 11 CCR section 7050(a)(5), a service provider may rely on the same purpose. The record’s legal-claims retention should be scoped to those provisions.

Is a model provider’s no-training commitment sufficient? The commitment binds the model provider, subject to the limitation of liability in the provider’s terms, and covers only the model provider. It says nothing about whether the customer’s own platform trains on, fine-tunes with or logs the data. The common gap in 2026 is a company that holds a no-training commitment from its foundation-model vendor while its own data processing addendum says nothing about what the company does with the same data.

Does this analysis apply to a company operating only in the United States? Yes. Section 3111 applies to any developer of a generative AI system made publicly available to Californians, wherever the developer is located. The CCPA applies to a business that does business in California and meets a threshold in section 1798.140(d), and its service-provider terms apply when the vendor processes personal information for that business. The European material applies to U.S. companies with European customers or European data subjects in their inputs.


I serve as outside product, commercial, and IP counsel to technology companies across the product lifecycle: platform and API terms, data processing and vendor agreements, privacy and data protection, and AI governance, including the acquisition, provenance and training-data questions this post raises. If these issues touch your product, get in touch.

This article is general information, not legal advice, and should not be relied on as such; it reflects the law as of September 2026, and reading it does not create an attorney-client relationship. Liam J Penwell is the professional name of Jason Penwell, Esq. (State Bar of California No. 339157), licensed in California. Penwell Law · Burlingame, California · Contact.

← Back to all posts

Subscribe

New posts, in your inbox.

About one email a week, whenever something new goes up: AI governance, privacy, and IP & content-rights commentary. No spam; unsubscribe in one click.

We’ll only use your email to send new posts. Unsubscribe anytime. See our Privacy Policy.