serp.fast

The provenance questions your web data vendor must answer before August 2

From 2 August 2026 the EU AI Office can enforce Article 53 training-data transparency. What that changes when you choose a web data or extraction vendor.

Nathan Kessler

Written by Nathan Kessler

Last updated: 8 min read

The provenance questions your web data vendor must answer before August 2

On 2 August 2026 the European Commission's AI Office gains the power to verify compliance with the EU AI Act's general-purpose AI obligations and to order corrective measures. The Commission's own FAQ on the training-content template is blunt about the consequence: publishing a summary of training content is mandatory, failure to do so "can lead to enforcement actions by the AI Office as of 2 August 2026", and non-compliance can carry fines of up to 3% of worldwide annual turnover or EUR 15 million, whichever is higher.

The obligation itself is not new. It has applied since 2 August 2025. What changes next week is that someone can act on it. And the second-order effect is the one worth planning around: the origin of a corpus stops being an internal engineering detail and becomes a document that a regulator, a customer's procurement team, or a publisher's counsel can ask to see. A vendor that cannot tell you where a given record came from was a reasonable cost saving in 2024. From August it is a gap in a paper trail.

What Article 53(1)(d) actually asks for

Article 53(1)(d) requires providers of general-purpose AI models placed on the EU market to publish a "sufficiently detailed summary" of the content used to train the model, following a template supplied by the AI Office, and explicitly including material protected by copyright. The Commission has said the template is mandatory and is the sole guidance for these summaries. Providers must refresh them at least every six months or whenever training data materially changes, and models already on the market before 2 August 2025 have until 2 August 2027.

Two details shape how much work this is. First, the obligation catches open-weight and free-licence releases too, not only commercial models. Second, per WilmerHale's reading of the regime, the Commission has said it will not run content-level audits of its own; enforcement is expected to follow complaints and "qualified alerts" from the Act's scientific panel. That matters for how the pressure will actually arrive. It will not arrive as a spot check of your S3 bucket. It will arrive as a publisher, a competitor, or a customer pointing at a specific source and asking whether it is in there.

Why this reaches buyers who are not model providers

Most people reading this are not placing a general-purpose model on the EU market, and the Act's obligation does not attach to them. The reach is commercial rather than direct, and it works in two ways.

If you further train or fine-tune an existing general-purpose model and release the result, the Commission's guidance already contemplates a separate summary covering the data used for that modification, referencing the original model's summary. That is a specific set of facts about specific fine-tuning, and whether it makes you a provider is a question for counsel, not a blog post. The point is that the boundary sits closer to normal product work than most teams assume.

The second route needs no legal theory at all. The obligation lands on a small number of model providers who each have thousands of suppliers, and those providers will push documentation requirements down the chain, because a summary they cannot substantiate is their exposure, not yours. Enterprise procurement does the same thing from the other side. Anyone who has filled in a security questionnaire knows how quickly an obligation on one party becomes a form everyone in the supply chain has to complete. The territorial question is similarly weak protection: the Act follows the market, so a US company selling into the EU is inside it, and there is no version of this where a vendor's Delaware incorporation answers the question.

Common Crawl is now a dependency with deletion pressure on it

In June 2026, Press Gazette reported that Digital Content Next, the US publisher trade body, sent Common Crawl a cease-and-desist letter through counsel. According to that reporting, the letter demanded Common Crawl immediately stop scraping, retaining or sharing copyrighted, paywalled and subscriber-only content, and remove publisher material already in its datasets. It also alleged that Common Crawl had made inaccurate or misleading statements about honouring opt-outs and removals. Press Gazette noted that more than 900 news sites appear in Common Crawl's opt-out registry.

Set the merits aside, because they are not settled. No lawsuit has been filed, this is a demand letter, and Common Crawl's executive director has previously denied misleading publishers about removals. What is relevant to a buyer is the shape of the risk rather than the outcome. A corpus under active deletion pressure is not a stable dependency for a retrieval index or an evaluation set. If snapshots are pruned, the artefact you built against changes underneath you, and any eval baseline computed from it stops being reproducible. Free was always the wrong reason to choose Common Crawl. Coverage and reproducibility were the good reasons, and one of them is now in question. Our guide to sourcing web data for LLM training covers where the alternatives sit.

Machine-readable terms are arriving faster than the standard that would unify them

The provenance question used to have a simple technical answer: read robots.txt, obey allow or deny, move on. That is no longer sufficient, because publishers have started expressing terms rather than a single bit.

RSL, published alongside robots.txt as a machine-readable licence, is the most visible of these. Its steering committee announced the RSL 1.0 specification in December 2025 and claims endorsement from more than 1,500 media organisations, brands and technology companies, naming Yahoo, Ziff Davis, O'Reilly Media, Cloudflare and Creative Commons among contributors and supporters. Treat those counts as what they are: figures from the standard's own promotion, not an audited adoption study. What matters more than the count is that a fetcher reading only allow-or-deny now discards terms the publisher deliberately put in machine-readable form, and leaves behind a record that is easy to point at afterwards.

The neutral path is the IETF's AIPREF working group, which is standardising a vocabulary for expressing AI usage preferences plus a mechanism for attaching those preferences to content through well-known URIs and HTTP response headers. Its datatracker milestones target 31 August 2026 for sending both the vocabulary and attachment specifications to the IESG. That is a target on a charter, not a shipped RFC, and the drafts are still drafts. Anyone telling you there is a settled standard to comply with is ahead of the facts.

Worth separating from all of this: llms.txt is not a provenance control. Rankability's June 2026 measurement of the Tranco top 1,000 domains found 8.7% publishing an llms.txt file, rising to 15.8% among the 549 domains that could actually serve a file at their root. It is a discovery and formatting convenience that a site publishes about itself. It says nothing about whether your vendor was entitled to fetch the page.

The five questions to send a vendor

Send these to any extraction, index or dataset supplier you depend on. The answers are more informative than the marketing.

  1. Does your corpus include Common Crawl snapshots, and which ones? A yes is not disqualifying. An "I'm not sure" is.
  2. Do you honour robots.txt, and do you parse machine-readable licence terms where they exist? Ask specifically whether the crawler reads anything beyond allow and deny.
  3. Can you produce per-source provenance for a given record? The test is a single result: give them a URL from an API response and ask when it was fetched, by what agent, and under what terms.
  4. What is your deletion process when a publisher demands removal? Ask how long it takes and whether removal propagates to derived artefacts such as embeddings and caches.
  5. What do you retain, and for how long? Retention is what makes questions three and four answerable a year later.

A vendor who answers all five in writing has effectively done part of your compliance work. A vendor who treats the questions as unusual has told you where their corpus came from without saying it. This slots into the broader checklist in our vendor due diligence guide.

The directory splits three ways on this

Exposure tracks how a vendor obtains its data, not how good the product is.

Licensed and first-party indexes answer origin questions structurally. Linkup licenses publisher content rather than scraping it, which turns provenance into a contract you can point at. Brave Search API and Ceramic run their own crawls and indexes, so the chain of custody has one owner. Webz.io sells structured feeds under commercial terms. Owning the index is what these vendors sell, and 2026 is the year that starts showing up in a procurement questionnaire rather than only in their marketing.

Open corpora are the exposed end. Common Crawl's value was always that everyone could reproduce everyone else's work against it. Deletion pressure attacks exactly that property.

Fetch-on-demand extraction sits in between and is better positioned than most buyers realise. Firecrawl, Zyte and Diffbot fetch a URL you asked for, at a time you can log, and return the result to you. There is no historical corpus of unknown origin in the middle. The provenance record is a request log, which is a solvable engineering problem rather than an archaeology project. The catch is that the record only exists if someone keeps it, which is a question about your pipeline as much as theirs.

What this means if you build on web data

Provenance is now a procurement artefact, not an engineering nicety. From 2 August the AI Office can act, and the pressure moves down the supply chain from the model providers who carry the obligation directly.

Treat "we used the open corpus" as a dependency risk, not a cost saving. A dataset facing removal demands cannot underwrite a reproducible evaluation, whatever happens to the underlying legal claim.

A compliant fetcher reads terms, not a single bit. Machine-readable licences are shipping now and the IETF vocabulary is targeted for late August 2026, so build the fetch layer to record what terms it saw rather than just whether it was allowed.

Ask the five questions in writing, before renewal. The answers are cheap to obtain today and expensive to reconstruct once someone else is asking you for them.

The uncomfortable version of all this: for most teams the build-versus-buy calculation has been about cost and success rates. A documentation burden is a new term in that equation, and it is one where a vendor with a real answer is worth paying for. The same dynamic drove the metered access layer we wrote about in June: that layer asked machines to identify themselves and pay, and Article 53 asks them to keep the receipts.

Share:

Tags:

  • #legal
  • #market-analysis
  • #data-quality
  • #llm-training
  • #api-selection