أصدرت وزارة العدل الأمريكية تدخلاً قانونياً هاماً في حرب حقوق الطبع والنشر المستمرة بين عمالقة الإعلام ومطوري الذكاء الاصطناعي. ومن خلال الدفع بأن تدريب نماذج اللغات الكبيرة (LLMs) على بيانات محمية بحقوق الطبع والنشر يشكل "استخداماً عادلاً"، وفرت وزارة العدل درعاً قانونياً ضخماً لشركات مثل OpenAI وMicrosoft.
حجة وزارة العدل: التدريب مقابل المخرجات
في قلب الدعوى القضائية الموحدة التي تشمل صحيفة The New York Times، يوجد تمييز جوهري بين كيفية تعلم الذكاء الاصطناعي وما ينتجه. تزعم صحيفة The New York Times أن الملايين من مقالاتها استُخدمت دون إذن لتدريب نماذج مثل GPT-4، مما تسبب في أضرار بمليارات الدولارات وأدى إلى إنشاء منتجات تنافس الصحيفة بشكل مباشر.
ومع ذلك، تجادل المذكرة الأخيرة لوزارة العدل بأن انتهاك حقوق الطبع والنشر يحدث عند نقطة المخرجات، وليس عند نقطة الاستيعاب. وتدفع الوزارة بأنه بينما يتم نسخ أعمال كاملة خلال مرحلة التدريب، إلا أن هذه النسخ لا تُتاح أبداً للجمهور. علاوة على ذلك، تؤكد وزارة العدل أن مخرجات نماذج مثل GPT-4 "غالباً ما تفتقر، إن لم يكن دائماً، إلى التشابه الجوهري" مع المادة المصدرية الأصلية. ومن خلال الفصل بين عملية التدريب والمحتوى المُنشأ، تحاول وزارة العدل فك الارتباط بين فعل تعلم الآلة والمفهوم القانوني للإحلال في السوق.
"تشبيه همنغواي" والإبداع البشري
لجعل مفهوم تعلم الآلة ملموساً، استشهدت وزارة العدل بتشبيه أدبي يتعلق بالكاتبة جوان ديديون. وأشارت المذكرة إلى أنه عندما كانت مراهقة، كانت ديديون تنسخ قصص إرنست همنغواي لتحليل بنية جمله وتعلم حرفته. وتجادل وزارة العدل بأنه إذا عامل القانون تدريب الآلة كأنه انتهاك، فقد تواجه كاتبة مثل ديديون مسؤولية قانونية نظرياً في كل مرة تنشر فيها عملاً جديداً، حيث ستكون عملية تعلمها مرتبطة ارتباطاً وثيقاً بكتابتها اللاحقة.
وتجادل الوزارة أيضاً بأن فرض مسؤولية صارمة على تدريب الذكاء الاصطناعي من شأنه، ويا للمفارقة، أن يخنق الإبداع ذاته الذي يهدف قانون حقوق الطبع والنشر إلى حمايته. ومع استخدام البشر المتزايد لنماذج LLMs لصياغة وتحرير الأعمال الأصلية، تشير وزارة العدل إلى أن جعل التدريب غير مسموح به دون مخططات ترخيص ضخمة من شأنه أن يشل الابتكار القائم على الذكاء الاصطناعي.
تحدي مكتب حقوق الطبع والنشر الأمريكي
يتعارض موقف وزارة العدل بشكل مباشر مع النتائج الأخيرة الصادرة عن مكتب حقوق الطبع والنشر الأمريكي. فقد جادلت المسجلة السابقة شيرا بيرلموتر سابقاً ضد دفاع "الاستخدام العادل" الشامل، مشيرة إلى أن الذكاء الاصطناعي يعمل بنطاق وسرعة يتجاوزان القدرات البشرية بكثير، وغالباً ما ينشئ منتجات تجارية تنافس منشئي المحتوى الأصليين بشكل مباشر.
وقد اتخذت وزارة العدل موقفاً هجومياً ضد هذا التقييم، صرحت فيه بأن تقرير بيرلموتر لا يحمل أي سلطة قانونية ملزمة ويفشل في مراعاة السوابق القضائية الراسخة المتعلقة بالتحليل لكل حالة على حدة. وتحذر الوزارة من أن فرض مسؤولية واسعة النطاق سيفرض فعلياً نظام ترخيص يجعل تطوير نماذج الذكاء الاصطناعي المتقدمة مستحيلاً من الناحية القانونية والمالية.
النقاط الرئيسية المستخلصة
- الفصل بين التدريب والمخرجات: تجادل وزارة العدل بأن نسخ النصوص للتدريب لا يشكل انتهاكاً إذا كانت مخرجات النموذج الناتجة لا تظهر "تشابهاً جوهرياً" مع الأعمال الأصلية.
- حماية الابتكار: تحذر الوزارة من أن اشتراط الحصول على تراخيص لجميع بيانات التدريب من شأنه أن يخنق الإبداع ويمنع تطوير نماذج LLMs التي تساعد في إنشاء المحتوى البشري.
- مؤشر قانوني: يخلق هذا التدخل صراعاً عالي المخاطر بين وزارة العدل ومكتب حقوق الطبع والنشر الأمريكي، مما يمهد الطريق لقرار قضائي حاسم بشأن مستقبل الذكاء الاصطناعي التوليدي.
ARTICLE: قدمت وزارة العدل الأمريكية مذكرة "صديق المحكمة" (amicus brief) في دعوى حقوق الطبع والنشر الخاصة بصحيفة New York Times، مجادلة بأن نسخ النصوص المحمية لتدريب نماذج اللغات الكبيرة (LLMs) يعد استخداماً عادلاً. وتمنح هذه المذكرة شركات مثل OpenAI وMicrosoft درعاً قانونياً قد يبعدها عن مجمع التعويضات الذي تقدر الصحيفة أنه سيصل إلى مليارات الدولارات.
لماذا تكتسب هذه القضية أهمية الآن
تجمع الدعوى القضائية عدة ادعاءات بأن مطوري الذكاء الاصطناعي قد غذوا نماذجهم بملايين المقالات الصحفية دون إذن، ثم طرحوا منتجات تنافس تقارير صحيفة Times نفسها بشكل مباشر. إذا رفضت المحكمة حجة وزارة العدل بشأن الاستخدام العادل، فقد تواجه شركات الذكاء الاصطناعي فواتير ترخيص ضخمة أو تُجبر على وقف تطوير الجيل القادم من الوكلاء الحواريين.
الحجة الأساسية لوزارة العدل
The brief splits the AI workflow into two distinct stages. First, the model ingests large corpora of text; second, it generates responses to user prompts. The department says infringement only arises when a copyrighted work is reproduced in a way that the public can access. Because the training copies never leave the developer’s servers, the act of ingestion does not meet that threshold.
The DOJ also leans on the “substantial similarity” test, a long-standing copyright standard. It contends that the text output by models such as GPT-4 “often if not always” differs enough from any single source that a plaintiff cannot prove the required similarity. In short, the brief argues that the legal focus should be on the final output, not the internal learning process.
A literary analogy to illustrate the point
To make the technical argument more relatable, the filing invokes a story about a teenage writer who copied Ernest Hemingway’s stories to study his style. The DOJ says that if the law treated that learning exercise as infringement, the writer could be sued each time she published a new piece, even though her work was original. The department warns that extending the same logic to AI would “paralyze the very creativity copyright law was designed to protect.”
The stakes for AI developers and content creators
- Innovation risk: Requiring licenses for every piece of text used in training could raise costs so high that building state-of-the-art models becomes financially untenable.
- Creative workflow: Human writers increasingly rely on LLMs for drafting, editing, and brainstorming. If the training process were deemed illegal, those tools could disappear, reshaping how content is produced across industries.
- Market impact: The newspaper industry argues that AI models that can answer questions or generate news-like text erode the value of original reporting, potentially siphoning advertising revenue and subscriptions.
The Copyright Office’s counterpoint
A recent report from the U.S. Copyright Office, authored under former Register Shira Perlmutter, pushed back against a blanket fair-use defense. The office highlighted three factors that set AI apart from human learning: the sheer volume of material copied, the speed at which it is processed, and the fact that the resulting models can be packaged and sold as commercial products that directly compete with the source creators.
The DOJ rebuts those points, noting that the report does not carry binding legal authority and that existing case law already requires a fact-by-fact analysis of fair use. It warns that imposing “broad liability” would effectively force the industry into a licensing regime that “would render the development of advanced AI models legally and financially impossible.”
What could happen next
The district court handling the Times case will have to weigh the DOJ’s fair-use position against the Copyright Office’s findings. A ruling in favor of the DOJ would set a precedent that training data can be harvested without explicit permission, provided the model’s outputs stay clear of substantial similarity. A decision siding with the newspaper could trigger a wave of licensing negotiations, potentially reshaping the economics of AI research.
Bottom line
The DOJ’s filing turns the question of whether AI training is fair use into a high-stakes courtroom battle. The outcome will determine whether developers can continue to build powerful language models on existing text or must seek costly licenses for every piece of material they ingest. The decision will ripple through the tech sector, the media industry, and the broader creative economy.
