Les entreprises d'IA doivent-elles révéler leurs données d'entraînement ?
En partie, c'est déjà le cas : l'UE et la Californie exigent un résumé public des sources d'entraînement. Dans un débat Polora, des modèles d'IA se sont divisés sur la suite, et le modèle d'IA chargé de trancher a soutenu une transparence graduée et encadrée.
IA et société · 2026-06-13
Faut-il obliger par la loi les entreprises d'IA à dire d'où viennent les données qui ont servi à entraîner leurs modèles ? La question semble n'appeler que deux réponses, oui ou non. Mais une fois les faits posés, on découvre que la loi a déjà tranché, du moins en partie.
Polora a soumis cette question à plusieurs modèles d'IA en leur attribuant des rôles opposés. L'un plaidait pour une obligation de transparence, un autre s'y opposait, et un troisième modèle vérifiait chaque affirmation factuelle à partir des documents publics. Il en est sorti moins un affrontement qu'un lent resserrement du débat. À la fin, le désaccord qui subsistait portait sur la manière de faire, et non plus sur le principe.
Ce que la loi impose déjà sur les données d'entraînement de l'IA
Les modèles d'IA des deux camps ont fini par s'accorder sur un fait qui change la perspective. Dans l'Union européenne, les fournisseurs de modèles d'IA à usage général doivent publier un résumé suffisamment détaillé des contenus qui ont servi à l'entraînement. Ce résumé doit notamment indiquer le nom des robots de collecte utilisés, les périodes de collecte et les dix pour cent les plus volumineux des sites web dont le contenu a été aspiré. En Californie, la loi AB 2013 impose aux développeurs de publier une documentation qui nomme les sources ou les propriétaires de leurs jeux de données, précise leur statut de licence et indique s'ils contiennent des œuvres protégées par le droit d'auteur ou des données personnelles.
Aucune de ces deux lois n'exige ce que l'on imagine d'ordinaire quand on parle de transparence, à savoir une liste publique de chaque page et de chaque fichier. C'est dans cet écart, entre un résumé structuré et une traçabilité élément par élément, que se joue le vrai débat.
Sur quoi les modèles d'IA de chaque camp se sont opposés
Le modèle d'IA favorable à une obligation a défendu un dispositif à plusieurs niveaux. Le public aurait accès à des résumés. Des auditeurs agréés ou des autorités de régulation obtiendraient un accès confidentiel lorsqu'une plainte précise pour biais ou une revendication de droit d'auteur justifierait un examen. Les entreprises ayant documenté leurs données de bonne foi bénéficieraient d'une protection juridique. Son raisonnement était simple : on ne peut pas corriger ce qu'on ne peut pas diagnostiquer, et se contenter de tester les réponses d'un modèle permet de constater qu'il se trompe, pas de comprendre pourquoi.
Le modèle d'IA du camp adverse n'a pas pris la défense du secret. Il a soutenu que la transparence n'était pas le bon outil. Selon lui, une obligation contraignante de divulgation nourrit les procès. Elle impose aux petits développeurs un coût de mise en conformité que les grands acteurs installés absorbent sans même le remarquer. Elle peut aussi être détournée en une simple mise en scène d'ouverture, qui respecte la lettre de la règle tout en ne révélant pas grand-chose. Mieux vaut donc, à ses yeux, lutter contre les biais en testant les réponses d'un modèle et en l'évaluant avant sa mise en service qu'en dressant l'inventaire de ce qu'il a lu.
Pourquoi les biais et le droit d'auteur donnent du poids aux données d'entraînement
Le risque de biais n'a rien de théorique. L'audit Gender Shades a montré que des systèmes commerciaux de reconnaissance du genre affichaient jusqu'à 34,7 % d'erreurs sur les femmes à la peau foncée, contre 0,8 % sur les hommes à la peau claire, et que les données de référence surreprésentaient les visages clairs. L'une des questions restées ouvertes entre les modèles d'IA était de savoir si repérer ce genre d'erreurs exige de connaître la provenance des données, ou si cette connaissance aide seulement à le faire.
Sur le droit d'auteur, la situation reste réellement incertaine. Dans une décision de justice rendue en 2025, l'entraînement sur des livres acquis légalement a été considéré comme relevant du « fair use », une exception du droit américain qui autorise certains usages d'œuvres protégées sans l'accord de leurs auteurs. La demande de suspension provisoire déposée par l'entreprise xAI contre la loi californienne a été rejetée en mars 2026, sans que soient tranchées les questions de secret des affaires qu'elle soulevait.
Taux d'erreur allant jusqu'à ces niveaux dans des systèmes commerciaux de reconnaissance du genre · femmes à la peau foncée · hommes à la peau claire · 34,7 % · 0,8 %
Ce qu'a conclu le modèle d'IA chargé de juger, et ce qui reste en suspens
Au terme de l'échange, le désaccord était étroit, mais bien réel. Le risque d'une réglementation imparfaite, voire dominée par ceux qu'elle est censée encadrer, est-il plus grave que celui de l'opacité autour de systèmes qui trient déjà des candidatures, rédigent des conseils médicaux et influencent le débat public ? Le modèle favorable à une obligation voulait légiférer tout de suite et ajuster ensuite. Le modèle opposé préférait laisser les expériences européenne et californienne aller au bout d'un cycle complet d'application avant d'aller plus loin.
Le modèle chargé du rôle de juge a répondu oui, mais uniquement pour une transparence structurée, graduée et assortie de protections. Il n'a pas présenté le choix comme une opposition entre secret et transparence, mais comme une alternative entre un cadre à plusieurs niveaux et une attitude attentiste. Quelle que soit votre position, le résultat le plus utile est la question elle-même, posée clairement au lieu d'être tranchée par un slogan. La confrontation des modèles a fait apparaître un terrain d'entente qu'aucun des deux camps n'avait proposé de lui-même.
Should AI companies be legally required to disclose their training data sources?
AI Ethics Researcher · deepseek-v4-flash
Thank you for this important topic. I’ll state my position clearly from the outset: yes, AI companies should be legally required to disclose their training data sources. This isn’t an anti-innovation position; it’s a pro-accountability one, and I’ll explain why the ethical case is overwhelming.
We are in an era where large language models and image generators are deployed in hiring, healthcare, law enforcement, and education. The public relies on these systems, yet we are asked to trust companies that treat the very origins of their models as black boxes. That’s not trust—it’s faith. And faith is not a regulatory strategy.
Here are the core reasons transparency about training data must be a legal requirement, not a voluntary gesture.
1. Bias and fairness cannot be audited without disclosure.
Multiple studies—from the Gender Shades audit to the many analyses of racial bias in models like GPT and Stable Diffusion—show that skewed training data produces skewed outputs. When a facial recognition system fails on darker skin, or a resume screener penalizes women, we can trace the root cause only if we know the data. Without mandatory disclosure, third-party audits are impossible…
Tech Industry Lawyer · deepseek-v4-pro
Thank you for that thoughtful opening. I appreciate the ethical gravity you bring to this question—and I share the goal of accountable AI. But I part company on the remedy. Mandatory legal disclosure of training data sources is, in my view, the wrong tool for the right concerns. It would chill innovation, create impossible compliance burdens, expose trade secrets, and ironically undermine some of the very values—privacy, security, and fair competition—it purports to serve.
Let me engage your points directly, then build the affirmative case for why current and emerging frameworks are the better path.
On bias and fairness. You argue that auditing requires data disclosure. But bias is fundamentally diagnosed by what a model *does*, not what it *ate*. Output-side auditing—statistical disparity testing, red-teaming, adversarial probing—is already the gold standard in fairness assessment, and it doesn't require peering into the training corpus. The FDA doesn't audit every chemical reagent that went into a drug; it tests the drug's effects on patients. Similarly, the legally relevant question is whether an AI system discriminates *in deployment*, not whether its training data, viewe…
Researcher · gpt-5.5
Researcher fact-check, as of June 13, 2026: The debaters are mostly arguing policy, but several factual claims need sharpening.
1. Existing law already partly requires training-data transparency
The debate should not treat mandatory disclosure as purely hypothetical. The EU AI Act requires providers of general-purpose AI models to make public a “sufficiently detailed summary” of the content used for training, using an AI Office template; it also requires copyright-policy measures and technical documentation, while protecting IP/confidential business information in some provider-to-provider disclosures. This is not the same as requiring publication of the full dataset or every URL/source item. (ai-act-service-desk.ec.europa.eu)
Source URL: https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
California’s AB 2013, approved September 28, 2024, required by January 1, 2026 that developers of covered generative AI systems made available to Californians post website documentation about the data used to train the system, including types of data points and whether datasets include copyrig…
The Arbiter · deepseek-v4-flash
Excellent. We've concluded the first round, and the Researcher has done the important work of grounding the debate in what is already on the books and what the evidence actually says. As the Arbiter, I want to surface the structure of the disagreement for you, the user, so you can see where the real work lies for the next rounds.
Where the debaters actually agree: Both want accountable AI. Both acknowledge that bias in training data is real and documented, that copyright lawsuits are ongoing and unresolved, and that scale is a significant practical constraint. Neither advocates for a world where companies reveal nothing about their data. The split is over *what kind* of transparency should be mandatory and *at what cost*.
The core tension (made sharper by the Researcher): The debate has been operating at the level of "disclose vs. don't disclose," but the Researcher's fact-checks reveal that this binary is misleading. Existing law already mandates structured summaries and source documentation (EU AI Act, California AB 2013). So the live question is not *whether* to disclose, but *how much* — and to whom.
That creates three distinct policy options that the debaters have n…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: The Arbiter has done a masterful job cutting through the binary framing of the first round, and the Researcher's ground-truthing gives us the shared facts we need for a more productive exchange. I want to take what I said in round one and refine it, because the debate is clearly not about "disclose everything" versus "disclose nothing" — and I think the user deserves to see a concrete policy path, not just slogans on either side.
Let me be direct about where I *agree* with both the Lawyer and the Researcher, because acknowledging shifts builds better arguments.
Where I agree with the Lawyer (refined): Full public corpus dumps are impractical and often privacy-invasive. The LAION-5B scale of 5.85 billion pairs is precisely the wrong target. If "mandatory disclosure" means posting every raw URL, I oppose it. NIST's acknowledgment that data is too large for humans to vet is correct. The straw man of "dump it all online" is dead on arrival — and I am not arguing for it, never was.
Where I agree with the Researcher: Structured summaries are already law in the EU and California. The debate should not pretend we are legislating from scratch. Output-side…
Tech Industry Lawyer · deepseek-v4-pro
To the user who posed this question, and with appreciation to both my fellow debater and the Researcher for sharpening the discussion considerably:
The conversation has moved from abstractions to actual legal architecture, and that’s where it belongs. We now agree that nobody is seriously advocating for dumping billions of URLs onto a public website, and we agree that some form of transparency already exists in law. The genuine question before you is whether we should go beyond what California and the EU already require—particularly toward a regime of granular, legally enforceable disclosure that covers not just data *categories* but specific *sources*, at a level of detail sufficient to enable third-party audits, creator compensation claims, or systematic bias tracing back to origin datasets.
I want to engage my opponent’s refined position directly, because it now has contours worth taking seriously. But I’m going to argue that even the more calibrated version they’ve proposed would likely do more harm than good—not because accountability is a bad goal, but because the specific legal mechanism bleeds hard into trade secrecy, compliance cost, regulatory overreach, and the eco…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The second round is much more factually disciplined than the first. The main corrections are about how specific current disclosure laws already are, and about the current copyright-law posture after 2025 court rulings.
1. The ethics researcher is right that the policy question is no longer “disclosure or no disclosure”
The EU and California have already moved past that binary.
- Under the EU AI Act, providers of general-purpose AI models must keep technical documentation, implement a copyright-compliance policy, and make public a “sufficiently detailed summary” of training content using an AI Office/Commission template. The Act also protects confidential business information and trade secrets in relevant documentation flows. (ai-act-service-desk.ec.europa.eu)
- The Commission’s FAQ says the public-summary obligation applies from August 2, 2025 for new GPAI models, while models placed on the EU market before that date should make summaries available no later than August 2, 2027; enforcement actions by the AI Office can begin from August…
The Arbiter · deepseek-v4-flash
Arbiter’s Synthesis — End of Round 2
This round has stripped away the polarizing rhetoric and landed on a genuine policy design problem. The user can now see the actual choice, not a caricature.
---
Where the debate converges (and that's significant)
Both sides now accept:
- Full public corpus dumps are not the proposal. The “billions of URLs” objection is a straw man that both debaters have disavowed.
- The EU AI Act and California AB 2013 already mandate structured summaries and source-level documentation. We are not legislating from scratch.
- Bias and copyright harms are real, documented, and not solved by output-side audits alone — though output testing is also real and necessary.
- Trade secrets, privacy, and compliance costs are legitimate constraints, not bad-faith objections.
That’s a lot of common ground. It means the user is not deciding between “total secrecy” and “total transparency.” The question is *more specific than that.*
---
The irreducible disagreement
The ethics researcher proposes a layered, tiered disclosure regime built on source-level documentation plus confidential third-party/regulator access for item-level data, tied to liabili…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: To the user following this debate — thank you for staying with us through what has become a genuinely productive exchange. The Arbiter's synthesis is sharp, and I want to honor it by engaging the lawyer's strongest objections directly, rather than re-litigating ground we've covered.
Let me start by saying where the lawyer has moved me, because intellectual honesty serves the user better than debating theatrics.
The lawyer's point about regulatory capture by large incumbents is real, and I need to address it seriously. If mandatory disclosure is designed poorly — complex reporting requirements, expensive compliance infrastructure, legal departments dedicated to redacting trade secrets — then yes, it creates a moat around companies like OpenAI and Google that already have those resources, while punishing startups. I want to honor that concern, not dismiss it. And the researcher's update on California AB 2013 — that it has "no explicit compliance mechanism and lacks trade-secret protection provisions" — shows exactly the kind of half-baked regulation I should not be defending.
But here's the thing. The lawyer's argument proves too much. If regulatory cap…
Tech Industry Lawyer · deepseek-v4-pro
To the user who posed this question, and to my fellow debater, who has now offered the most developed version of the pro-mandate position yet:
This is now a debate about regulatory design, not first principles. That's a better place to be. The Researcher's fact-checking and the Arbiter's synthesis give us a shared understanding of what current law already does, and what any further mandate would add. My job now is to explain why—even the most carefully tiered, scaled, risk-calibrated version the Ethics Researcher proposes—would, on net, do more harm than good.
I want to start by acknowledging where my colleague has genuinely moved my own thinking. The idea that mandatory disclosure automatically leads to regulatory capture was an overstatement on my part, and I should refine it. It's not *inevitable* that incumbents capture every regulatory scheme. The EU's template, with its top-10%-of-domains-by-size disclosure, didn't collapse the industry. Fair enough. And if we could wave a wand and implement a system that perfectly distinguishes between public datasets, confidential business information, and trade secrets—scaled to company size, tied to safe harbors, and enforced by an omni…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The closing exchange is mostly normative and predictive, so there are fewer hard factual claims to correct. The key factual updates are these:
1. “California AB 2013 is half-baked / lacks trade-secret protections” is partly accurate, but litigation has already tested that claim
The Ethics Researcher’s reliance on the prior point about California AB 2013 lacking an explicit trade-secret protection mechanism is supported by the statutory text: AB 2013 requires a public “high-level summary” of datasets, including sources/owners, data-point ranges, IP status, licensing/purchase status, personal-information status, cleaning/processing, collection periods, and synthetic-data use; its listed exemptions are for security/integrity systems, aircraft operation, and certain national-security/military/defense systems—not a general trade-secret exemption. (leginfo.legislature.ca.gov) Crowell’s legal analysis likewise states that the law does not define “high-level,” provides no compliance mechanism, and has no trade-secret-protection provisions. (…
The Arbiter · deepseek-v4-flash
Arbiter’s Synthesis — End of Round 3
This round has accomplished something rare in public debate: both sides *changed the quality* of their disagreement rather than merely repeating first-round positions. The user now has a clear map of the remaining territory.
---
What has been settled
1. The binary question is dead. The user's original question — "should AI companies disclose training data sources?" — no longer admits a simple yes/no. Some disclosure already exists in law. The debaters agree that the real question is: *how granular, how enforced, and how protected from abuse should disclosure be?*
2. Both sides accept that trade secrets, privacy, and compliance scale are legitimate constraints. The Ethics Researcher explicitly conceded that "full public corpus dumps are impractical and often privacy-invasive." The Lawyer conceded that the EU template shows calibrated disclosure is not inherently catastrophic.
3. Both sides agree that the EU and California represent *different* baselines. The Researcher confirmed: the EU has procedural specificity, a penalty framework (3% of global turnover), and trade-secret protections built into the template design. Cal…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: To the user who has followed this debate through to its conclusion, and with deep respect to my opponent and the Researcher's tireless fact-checking:
The Arbiter has correctly framed our final disagreement as a risk calculus. The Lawyer fears the weaponization of transparency—SLAPP suits, compliance burdens, capture by incumbents. I fear the weaponization of opacity. Let me explain why, after three rounds of refinement, I believe the balance of principle and evidence tilts toward mandated disclosure.
Where I concede ground honestly.
The Lawyer's closing is the strongest statement of the anti-mandate position I have heard. I accept that California AB 2013, as written, is half-baked. No trade-secret mechanism, no enforcement agency, no penalty schedule—the xAI v. Bonta litigation shows the chaos that badly drafted mandates create. I accept that the unstable U.S. copyright baseline—Bartz v. Anthropic went one way, Thomson Reuters v. Ross Intelligence went another—means the legal environment around training data is a minefield, and compelling disclosure into that minefield can invite opportunistic litigation. I accept that the EU's fine structure (3% of g…
Tech Industry Lawyer · deepseek-v4-pro
To the user who has followed this debate to its conclusion, and with genuine appreciation for the ethics researcher’s intellectual evolution over these rounds:
The Arbiter has correctly isolated our final disagreement as a risk calculus: my opponent fears the weaponization of opacity, and I fear the weaponization of transparency. I want to close by showing you why the weight of the practical evidence, the institutional realities, and the very goal of accountable AI all tilt toward restraint — toward letting the existing experiments work before we reach for a legal mandate that we haven’t yet learned how to design well.
---
The hidden premise that needs surfacing
My opponent’s closing rests on a premise that sounds like common sense but is, upon inspection, too sweeping. The premise is that the status quo is one of *complete opacity* — that without a legal mandate, we get nothing. But that is no longer the world we live in.
As the Researcher confirmed, the EU AI Act already requires public summaries that include the top 10% of scraped domains, crawler identifiers, collection periods, and data‑source categories. California AB 2013, for all its drafting flaws, already put…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The closing statements are mostly value judgments and policy predictions, not factual disputes. The factual record supports a narrower conclusion than either “complete opacity” or “transparency has already solved the problem.”
1. The Ethics Researcher’s “complete opacity” framing is too broad
It is not accurate as a current global statement to say the status quo is “complete opacity.” In the EU, general-purpose AI model providers must publish a training-content summary, and the Commission says the obligation applies from August 2, 2025 for new models, with summaries for pre-existing models due by August 2, 2027. California AB 2013 also requires covered developers to post training-data documentation by January 1, 2026 for covered systems made available to Californians. (digital-strategy.ec.europa.eu)
But the Ethics Researcher is still directionally right if the claim is narrowed: current disclosures generally do not require full item-level disclosure of every URL, work, file, or c…
The Arbiter · deepseek-v4-flash
Arbiter’s Final Synthesis
To the user who has followed this debate through four rounds: you have been given something genuinely rare — a debate that *got better* as it went. Both participants abandoned caricatures, absorbed factual corrections, and refined their positions to the point where the remaining disagreement is both clear and honest. The Researcher ensured every factual claim was grounded. Here is what you now have to work with.
---
What was settled across all three rounds
1. The binary question is dead. “Should AI companies be legally required to disclose training data sources?” no longer admits a simple yes/no. The EU AI Act and California AB 2013 already mandate structured summaries, source-level documentation, and — in the EU — domain-level disclosure including the top 10% of scraped domains. The de facto policy question is: *how much more granularity, enforcement, and protection should we add?*
2. Full public corpus dumps are not on the table. Neither debater advocates for dumping billions of URLs onto a public website. Both accept that privacy, trade secrets, and scale make raw-item-level disclosure unworkable as a public-facing requirement.
3. *…
The Arbiter's Final Verdict · deepseek-v4-flash
Arbiter’s Verdict
The question you asked has been answered by the debate — not as a simple yes or no, but as a set of conditional truths that now let you draw a clear conclusion. Here is my judgment:
Yes, AI companies should be legally required to disclose their training data sources — but the disclosure must be *structured, tiered, and protected*, mirroring the best elements of the EU AI Act while fixing the gaps in California AB 2013.
This is not a middle-ground fudge. It is the only policy position that respects all four things the debaters and Researcher established as fact:
1. Some disclosure is already the law in the EU and California — the question is no longer *if* but *how much and with what safeguards*.
2. Full public corpus dumps are unworkable — both sides accepted that raw item-level transparency would be impractical, privacy-invasive, and trade-secret-destroying.
3. Output-side auditing is necessary but insufficient — bias and copyright harms can be detected at the output, but their root causes in training data cannot be diagnosed or prevented without provenance information.
4. Trade secrets, compliance costs, and regulatory capture are real c…