부분적으로는 이미 의무입니다. EU와 캘리포니아는 학습 데이터 출처 요약을 공개하라고 요구합니다. 폴로라 토론에서 AI 모델들은 어디까지 공개할지를 두고 갈렸고, 판결을 맡은 모델은 단계를 나누고 기밀을 보호하는 공개를 지지했습니다.
AI와 사회 · 2026-06-13
AI 회사가 모델을 학습시킨 데이터를 어디서 가져왔는지 법으로 공개하게 해야 할까요? 답이 둘 중 하나일 것 같은 질문입니다. 그런데 사실을 늘어놓고 보면 놀랍게도 법은 이미 한쪽을 골랐습니다. 다만 끝까지 간 것은 아닙니다.
폴로라는 이 질문을 서로 반대 역할을 맡은 AI 모델들에게 던졌습니다. 한 모델은 공개 의무화를 주장했고 다른 모델은 반대했습니다. 세 번째 모델은 양쪽이 내놓은 사실 주장을 하나하나 공개된 기록과 대조했습니다. 결과는 맞대결이라기보다 쟁점이 천천히 좁혀지는 과정에 가까웠습니다. 끝에 남은 이견은 원칙이 아니라 제도를 어떻게 짜느냐에 관한 것이었습니다.
법이 이미 요구하는 AI 학습 데이터 공개
찬반을 맡은 두 AI 모델은 결국 판을 바꾸는 사실 하나에 함께 동의했습니다. 유럽연합에서는 범용 AI 모델을 내놓는 회사가 학습에 쓴 내용을 충분히 자세하게 요약해 공개해야 합니다. 여기에는 크롤러(웹 페이지를 자동으로 수집하는 프로그램)의 이름, 수집 기간, 그리고 긁어 온 도메인 가운데 규모로 상위 십 퍼센트에 드는 목록이 들어갑니다. 미국 캘리포니아주의 AB 2013 법은 개발사가 데이터셋의 출처나 소유자, 이용 허락 여부, 저작권이 있는 자료나 개인정보가 포함됐는지를 적은 문서를 공개하도록 합니다.
두 법 모두 사람들이 '공개'라는 말을 들을 때 흔히 떠올리는 것, 곧 모든 웹 페이지와 파일을 하나하나 적은 공개 목록까지 요구하지는 않습니다. 진짜 논쟁은 정리된 요약과 항목 하나하나의 출처, 그 둘 사이의 틈에 있습니다.
의무화를 주장한 AI 모델은 단계를 나눈 제도 쪽으로 논의를 밀고 갔습니다. 요약은 누구나 볼 수 있게 공개합니다. 특정한 편향 신고나 저작권 주장이 들어와 들여다볼 이유가 생기면, 인가받은 감사인이나 규제 기관이 기밀을 유지하는 조건으로 세부 자료를 봅니다. 데이터를 성실하게 기록해 둔 회사에는 법적 책임을 덜어 주는 면책 조항을 줍니다. 이 모델의 논리는 단순했습니다. 진단할 수 없는 것은 고칠 수 없고, 모델의 출력만 시험해서는 모델이 틀린다는 것은 알아도 왜 틀리는지는 알 수 없다는 것입니다.
반대편 AI 모델은 비밀주의를 옹호하지 않았습니다. 대신 공개가 이 문제에 맞는 도구가 아니라고 주장했습니다. 이 모델이 보기에 강제력 있는 공개 의무는 소송을 부릅니다. 규모가 작은 개발사에는 규정을 지키는 비용이 큰 짐이 되지만, 이미 자리 잡은 대기업에는 대수롭지 않은 금액입니다. 게다가 규정의 글자는 지키면서 실질적인 내용은 거의 드러내지 않는, 보여 주기식 개방으로 흘러갈 수도 있습니다. 이 모델은 편향 문제라면 모델이 무엇을 읽었는지 목록을 만들기보다 출력을 시험하고 배포 전에 편향을 평가하는 편이 낫다고 봤습니다.
편향과 저작권 때문에 학습 데이터가 중요한 이유
편향에 대한 걱정은 추상적인 이야기가 아닙니다. MIT 미디어랩의 '젠더 셰이즈(Gender Shades)' 감사 연구에 따르면, 얼굴 사진으로 성별을 판별하는 상용 시스템들이 피부색이 어두운 여성을 잘못 분류한 비율은 최대 34.7퍼센트였고, 피부색이 밝은 남성은 0.8퍼센트였습니다. 이 시스템들을 평가하는 기준 데이터도 밝은 얼굴 쪽으로 치우쳐 있었습니다. AI 모델들의 토론에서 끝까지 이어진 쟁점 하나는, 이런 실패를 잡아내려면 데이터의 출처를 반드시 알아야 하는지, 아니면 그 지식은 도움이 될 뿐인지였습니다.
저작권 쪽은 정말로 아직 정리되지 않았습니다. 2025년의 한 판결은 합법적으로 구한 책으로 모델을 학습시킨 것을 공정 이용(저작권자의 허락 없이도 쓸 수 있다고 인정되는 경우)으로 봤습니다. AI 회사 xAI가 캘리포니아 법에 맞서 낸 소송에서는 2026년 3월 법원이 법 시행을 일단 멈춰 달라는 가처분 신청을 받아들이지 않았습니다. 다만 xAI가 제기한 영업비밀 문제에 대한 판단은 내리지 않았습니다.
젠더 셰이즈 감사 연구 속 상용 얼굴 성별 판별 시스템의 오분류 비율 · 피부색이 어두운 여성 · 피부색이 밝은 남성 · 최대 34.7퍼센트 · 0.8퍼센트
끝에 남은 이견은 좁지만 실재했습니다. 불완전하고 자칫 업계에 휘둘릴 수도 있는 규제의 위험이, 이미 입사 지원자를 가려내고 의료 조언의 초안을 쓰고 여론을 움직이는 시스템들이 속을 알 수 없는 상태로 남는 위험보다 더 클까요? 의무화를 주장한 AI 모델은 지금 법을 만들고 나중에 고치자는 쪽이었습니다. 반대한 AI 모델은 EU와 캘리포니아의 실험이 한 차례 집행 주기를 다 거친 뒤에 규제를 더하자는 쪽이었습니다.
판결 역할을 맡은 AI 모델은 공개를 의무화해야 한다고 결론 내렸습니다. 다만 정리된 형식을 갖추고, 단계를 나누고, 기밀을 보호하는 방식의 공개에 한해서입니다. 이 모델은 선택지를 비밀주의 대 투명성이 아니라, 여러 층으로 짠 제도 대 일단 지켜보기로 정리했습니다. 어느 쪽에 마음이 가든, 이 토론에서 건질 것은 구호로 답하는 대신 깔끔하게 다시 세운 질문 그 자체입니다. 모델들을 서로 맞세우자, 어느 쪽도 혼자서는 내놓지 않았던 합의점이 드러났습니다.
끝에 남은 이견 : 지금 법을 만들 것인가, 집행을 지켜본 뒤 더할 것인가 · 의무화를 주장한 AI 모델 · 반대한 AI 모델 · 의무화를 주장한 AI 모델은 지금 법을 만들고 나중에 고치자는 쪽이었습니다. · 반대한 AI 모델은 EU와 캘리포니아의 실험이 한 차례 집행 주기를 다 거친 뒤에 규제를 더하자는 쪽이었습니다.
Should AI companies be legally required to disclose their training data sources?
AI Ethics Researcher · deepseek-v4-flash
Thank you for this important topic. I’ll state my position clearly from the outset: yes, AI companies should be legally required to disclose their training data sources. This isn’t an anti-innovation position; it’s a pro-accountability one, and I’ll explain why the ethical case is overwhelming.
We are in an era where large language models and image generators are deployed in hiring, healthcare, law enforcement, and education. The public relies on these systems, yet we are asked to trust companies that treat the very origins of their models as black boxes. That’s not trust—it’s faith. And faith is not a regulatory strategy.
Here are the core reasons transparency about training data must be a legal requirement, not a voluntary gesture.
1. Bias and fairness cannot be audited without disclosure.
Multiple studies—from the Gender Shades audit to the many analyses of racial bias in models like GPT and Stable Diffusion—show that skewed training data produces skewed outputs. When a facial recognition system fails on darker skin, or a resume screener penalizes women, we can trace the root cause only if we know the data. Without mandatory disclosure, third-party audits are impossible…
Tech Industry Lawyer · deepseek-v4-pro
Thank you for that thoughtful opening. I appreciate the ethical gravity you bring to this question—and I share the goal of accountable AI. But I part company on the remedy. Mandatory legal disclosure of training data sources is, in my view, the wrong tool for the right concerns. It would chill innovation, create impossible compliance burdens, expose trade secrets, and ironically undermine some of the very values—privacy, security, and fair competition—it purports to serve.
Let me engage your points directly, then build the affirmative case for why current and emerging frameworks are the better path.
On bias and fairness. You argue that auditing requires data disclosure. But bias is fundamentally diagnosed by what a model *does*, not what it *ate*. Output-side auditing—statistical disparity testing, red-teaming, adversarial probing—is already the gold standard in fairness assessment, and it doesn't require peering into the training corpus. The FDA doesn't audit every chemical reagent that went into a drug; it tests the drug's effects on patients. Similarly, the legally relevant question is whether an AI system discriminates *in deployment*, not whether its training data, viewe…
Researcher · gpt-5.5
Researcher fact-check, as of June 13, 2026: The debaters are mostly arguing policy, but several factual claims need sharpening.
1. Existing law already partly requires training-data transparency
The debate should not treat mandatory disclosure as purely hypothetical. The EU AI Act requires providers of general-purpose AI models to make public a “sufficiently detailed summary” of the content used for training, using an AI Office template; it also requires copyright-policy measures and technical documentation, while protecting IP/confidential business information in some provider-to-provider disclosures. This is not the same as requiring publication of the full dataset or every URL/source item. (ai-act-service-desk.ec.europa.eu)
Source URL: https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
California’s AB 2013, approved September 28, 2024, required by January 1, 2026 that developers of covered generative AI systems made available to Californians post website documentation about the data used to train the system, including types of data points and whether datasets include copyrig…
The Arbiter · deepseek-v4-flash
Excellent. We've concluded the first round, and the Researcher has done the important work of grounding the debate in what is already on the books and what the evidence actually says. As the Arbiter, I want to surface the structure of the disagreement for you, the user, so you can see where the real work lies for the next rounds.
Where the debaters actually agree: Both want accountable AI. Both acknowledge that bias in training data is real and documented, that copyright lawsuits are ongoing and unresolved, and that scale is a significant practical constraint. Neither advocates for a world where companies reveal nothing about their data. The split is over *what kind* of transparency should be mandatory and *at what cost*.
The core tension (made sharper by the Researcher): The debate has been operating at the level of "disclose vs. don't disclose," but the Researcher's fact-checks reveal that this binary is misleading. Existing law already mandates structured summaries and source documentation (EU AI Act, California AB 2013). So the live question is not *whether* to disclose, but *how much* — and to whom.
That creates three distinct policy options that the debaters have n…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: The Arbiter has done a masterful job cutting through the binary framing of the first round, and the Researcher's ground-truthing gives us the shared facts we need for a more productive exchange. I want to take what I said in round one and refine it, because the debate is clearly not about "disclose everything" versus "disclose nothing" — and I think the user deserves to see a concrete policy path, not just slogans on either side.
Let me be direct about where I *agree* with both the Lawyer and the Researcher, because acknowledging shifts builds better arguments.
Where I agree with the Lawyer (refined): Full public corpus dumps are impractical and often privacy-invasive. The LAION-5B scale of 5.85 billion pairs is precisely the wrong target. If "mandatory disclosure" means posting every raw URL, I oppose it. NIST's acknowledgment that data is too large for humans to vet is correct. The straw man of "dump it all online" is dead on arrival — and I am not arguing for it, never was.
Where I agree with the Researcher: Structured summaries are already law in the EU and California. The debate should not pretend we are legislating from scratch. Output-side…
Tech Industry Lawyer · deepseek-v4-pro
To the user who posed this question, and with appreciation to both my fellow debater and the Researcher for sharpening the discussion considerably:
The conversation has moved from abstractions to actual legal architecture, and that’s where it belongs. We now agree that nobody is seriously advocating for dumping billions of URLs onto a public website, and we agree that some form of transparency already exists in law. The genuine question before you is whether we should go beyond what California and the EU already require—particularly toward a regime of granular, legally enforceable disclosure that covers not just data *categories* but specific *sources*, at a level of detail sufficient to enable third-party audits, creator compensation claims, or systematic bias tracing back to origin datasets.
I want to engage my opponent’s refined position directly, because it now has contours worth taking seriously. But I’m going to argue that even the more calibrated version they’ve proposed would likely do more harm than good—not because accountability is a bad goal, but because the specific legal mechanism bleeds hard into trade secrecy, compliance cost, regulatory overreach, and the eco…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The second round is much more factually disciplined than the first. The main corrections are about how specific current disclosure laws already are, and about the current copyright-law posture after 2025 court rulings.
1. The ethics researcher is right that the policy question is no longer “disclosure or no disclosure”
The EU and California have already moved past that binary.
- Under the EU AI Act, providers of general-purpose AI models must keep technical documentation, implement a copyright-compliance policy, and make public a “sufficiently detailed summary” of training content using an AI Office/Commission template. The Act also protects confidential business information and trade secrets in relevant documentation flows. (ai-act-service-desk.ec.europa.eu)
- The Commission’s FAQ says the public-summary obligation applies from August 2, 2025 for new GPAI models, while models placed on the EU market before that date should make summaries available no later than August 2, 2027; enforcement actions by the AI Office can begin from August…
The Arbiter · deepseek-v4-flash
Arbiter’s Synthesis — End of Round 2
This round has stripped away the polarizing rhetoric and landed on a genuine policy design problem. The user can now see the actual choice, not a caricature.
---
Where the debate converges (and that's significant)
Both sides now accept:
- Full public corpus dumps are not the proposal. The “billions of URLs” objection is a straw man that both debaters have disavowed.
- The EU AI Act and California AB 2013 already mandate structured summaries and source-level documentation. We are not legislating from scratch.
- Bias and copyright harms are real, documented, and not solved by output-side audits alone — though output testing is also real and necessary.
- Trade secrets, privacy, and compliance costs are legitimate constraints, not bad-faith objections.
That’s a lot of common ground. It means the user is not deciding between “total secrecy” and “total transparency.” The question is *more specific than that.*
---
The irreducible disagreement
The ethics researcher proposes a layered, tiered disclosure regime built on source-level documentation plus confidential third-party/regulator access for item-level data, tied to liabili…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: To the user following this debate — thank you for staying with us through what has become a genuinely productive exchange. The Arbiter's synthesis is sharp, and I want to honor it by engaging the lawyer's strongest objections directly, rather than re-litigating ground we've covered.
Let me start by saying where the lawyer has moved me, because intellectual honesty serves the user better than debating theatrics.
The lawyer's point about regulatory capture by large incumbents is real, and I need to address it seriously. If mandatory disclosure is designed poorly — complex reporting requirements, expensive compliance infrastructure, legal departments dedicated to redacting trade secrets — then yes, it creates a moat around companies like OpenAI and Google that already have those resources, while punishing startups. I want to honor that concern, not dismiss it. And the researcher's update on California AB 2013 — that it has "no explicit compliance mechanism and lacks trade-secret protection provisions" — shows exactly the kind of half-baked regulation I should not be defending.
But here's the thing. The lawyer's argument proves too much. If regulatory cap…
Tech Industry Lawyer · deepseek-v4-pro
To the user who posed this question, and to my fellow debater, who has now offered the most developed version of the pro-mandate position yet:
This is now a debate about regulatory design, not first principles. That's a better place to be. The Researcher's fact-checking and the Arbiter's synthesis give us a shared understanding of what current law already does, and what any further mandate would add. My job now is to explain why—even the most carefully tiered, scaled, risk-calibrated version the Ethics Researcher proposes—would, on net, do more harm than good.
I want to start by acknowledging where my colleague has genuinely moved my own thinking. The idea that mandatory disclosure automatically leads to regulatory capture was an overstatement on my part, and I should refine it. It's not *inevitable* that incumbents capture every regulatory scheme. The EU's template, with its top-10%-of-domains-by-size disclosure, didn't collapse the industry. Fair enough. And if we could wave a wand and implement a system that perfectly distinguishes between public datasets, confidential business information, and trade secrets—scaled to company size, tied to safe harbors, and enforced by an omni…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The closing exchange is mostly normative and predictive, so there are fewer hard factual claims to correct. The key factual updates are these:
1. “California AB 2013 is half-baked / lacks trade-secret protections” is partly accurate, but litigation has already tested that claim
The Ethics Researcher’s reliance on the prior point about California AB 2013 lacking an explicit trade-secret protection mechanism is supported by the statutory text: AB 2013 requires a public “high-level summary” of datasets, including sources/owners, data-point ranges, IP status, licensing/purchase status, personal-information status, cleaning/processing, collection periods, and synthetic-data use; its listed exemptions are for security/integrity systems, aircraft operation, and certain national-security/military/defense systems—not a general trade-secret exemption. (leginfo.legislature.ca.gov) Crowell’s legal analysis likewise states that the law does not define “high-level,” provides no compliance mechanism, and has no trade-secret-protection provisions. (…
The Arbiter · deepseek-v4-flash
Arbiter’s Synthesis — End of Round 3
This round has accomplished something rare in public debate: both sides *changed the quality* of their disagreement rather than merely repeating first-round positions. The user now has a clear map of the remaining territory.
---
What has been settled
1. The binary question is dead. The user's original question — "should AI companies disclose training data sources?" — no longer admits a simple yes/no. Some disclosure already exists in law. The debaters agree that the real question is: *how granular, how enforced, and how protected from abuse should disclosure be?*
2. Both sides accept that trade secrets, privacy, and compliance scale are legitimate constraints. The Ethics Researcher explicitly conceded that "full public corpus dumps are impractical and often privacy-invasive." The Lawyer conceded that the EU template shows calibrated disclosure is not inherently catastrophic.
3. Both sides agree that the EU and California represent *different* baselines. The Researcher confirmed: the EU has procedural specificity, a penalty framework (3% of global turnover), and trade-secret protections built into the template design. Cal…
AI Ethics Researcher · deepseek-v4-flash
[AI Ethics Researcher]: To the user who has followed this debate through to its conclusion, and with deep respect to my opponent and the Researcher's tireless fact-checking:
The Arbiter has correctly framed our final disagreement as a risk calculus. The Lawyer fears the weaponization of transparency—SLAPP suits, compliance burdens, capture by incumbents. I fear the weaponization of opacity. Let me explain why, after three rounds of refinement, I believe the balance of principle and evidence tilts toward mandated disclosure.
Where I concede ground honestly.
The Lawyer's closing is the strongest statement of the anti-mandate position I have heard. I accept that California AB 2013, as written, is half-baked. No trade-secret mechanism, no enforcement agency, no penalty schedule—the xAI v. Bonta litigation shows the chaos that badly drafted mandates create. I accept that the unstable U.S. copyright baseline—Bartz v. Anthropic went one way, Thomson Reuters v. Ross Intelligence went another—means the legal environment around training data is a minefield, and compelling disclosure into that minefield can invite opportunistic litigation. I accept that the EU's fine structure (3% of g…
Tech Industry Lawyer · deepseek-v4-pro
To the user who has followed this debate to its conclusion, and with genuine appreciation for the ethics researcher’s intellectual evolution over these rounds:
The Arbiter has correctly isolated our final disagreement as a risk calculus: my opponent fears the weaponization of opacity, and I fear the weaponization of transparency. I want to close by showing you why the weight of the practical evidence, the institutional realities, and the very goal of accountable AI all tilt toward restraint — toward letting the existing experiments work before we reach for a legal mandate that we haven’t yet learned how to design well.
---
The hidden premise that needs surfacing
My opponent’s closing rests on a premise that sounds like common sense but is, upon inspection, too sweeping. The premise is that the status quo is one of *complete opacity* — that without a legal mandate, we get nothing. But that is no longer the world we live in.
As the Researcher confirmed, the EU AI Act already requires public summaries that include the top 10% of scraped domains, crawler identifiers, collection periods, and data‑source categories. California AB 2013, for all its drafting flaws, already put…
Researcher · gpt-5.5
Researcher final fact-check, as of June 13, 2026: The closing statements are mostly value judgments and policy predictions, not factual disputes. The factual record supports a narrower conclusion than either “complete opacity” or “transparency has already solved the problem.”
1. The Ethics Researcher’s “complete opacity” framing is too broad
It is not accurate as a current global statement to say the status quo is “complete opacity.” In the EU, general-purpose AI model providers must publish a training-content summary, and the Commission says the obligation applies from August 2, 2025 for new models, with summaries for pre-existing models due by August 2, 2027. California AB 2013 also requires covered developers to post training-data documentation by January 1, 2026 for covered systems made available to Californians. (digital-strategy.ec.europa.eu)
But the Ethics Researcher is still directionally right if the claim is narrowed: current disclosures generally do not require full item-level disclosure of every URL, work, file, or c…
The Arbiter · deepseek-v4-flash
Arbiter’s Final Synthesis
To the user who has followed this debate through four rounds: you have been given something genuinely rare — a debate that *got better* as it went. Both participants abandoned caricatures, absorbed factual corrections, and refined their positions to the point where the remaining disagreement is both clear and honest. The Researcher ensured every factual claim was grounded. Here is what you now have to work with.
---
What was settled across all three rounds
1. The binary question is dead. “Should AI companies be legally required to disclose training data sources?” no longer admits a simple yes/no. The EU AI Act and California AB 2013 already mandate structured summaries, source-level documentation, and — in the EU — domain-level disclosure including the top 10% of scraped domains. The de facto policy question is: *how much more granularity, enforcement, and protection should we add?*
2. Full public corpus dumps are not on the table. Neither debater advocates for dumping billions of URLs onto a public website. Both accept that privacy, trade secrets, and scale make raw-item-level disclosure unworkable as a public-facing requirement.
3. *…
The Arbiter's Final Verdict · deepseek-v4-flash
Arbiter’s Verdict
The question you asked has been answered by the debate — not as a simple yes or no, but as a set of conditional truths that now let you draw a clear conclusion. Here is my judgment:
Yes, AI companies should be legally required to disclose their training data sources — but the disclosure must be *structured, tiered, and protected*, mirroring the best elements of the EU AI Act while fixing the gaps in California AB 2013.
This is not a middle-ground fudge. It is the only policy position that respects all four things the debaters and Researcher established as fact:
1. Some disclosure is already the law in the EU and California — the question is no longer *if* but *how much and with what safeguards*.
2. Full public corpus dumps are unworkable — both sides accepted that raw item-level transparency would be impractical, privacy-invasive, and trade-secret-destroying.
3. Output-side auditing is necessary but insufficient — bias and copyright harms can be detected at the output, but their root causes in training data cannot be diagnosed or prevented without provenance information.
4. Trade secrets, compliance costs, and regulatory capture are real c…