Close Menu
  • Home
  • Stock
  • Parenting
  • Personal
  • Fashion & Beauty
  • Finance & Business
  • Marketing
  • Health & Fitness
  • Tech & Gadgets
  • Travel & Adventure

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

What's Hot

4 longevity mistakes that catch up to women after menopause

agosto 24, 2026

‘Lust-free’ gyms say they’re safe spaces for Christian men

agosto 21, 2026

Experimental ‘exercise pill’ may reduce appetite and mimic physical activity

agosto 21, 2026
Facebook X (Twitter) Instagram
  • Home
  • Contact us
  • DMCA
  • Política de Privacidad
  • Publicidad en DD Noticias
  • Sobre Nosotros
  • Términos y Condiciones
Facebook X (Twitter) Instagram
DD Noticias: Tu fuente de inspiración diariaDD Noticias: Tu fuente de inspiración diaria
  • Home
  • Stock
  • Parenting
  • Personal
  • Fashion & Beauty
  • Finance & Business
  • Marketing
  • Health & Fitness
  • Tech & Gadgets
  • Travel & Adventure
DD Noticias: Tu fuente de inspiración diariaDD Noticias: Tu fuente de inspiración diaria
Home » OpenAI Trained AI Models on Copyrighted O’Reilly Media Books, Researchers Claim
Technology & Gadgets

OpenAI Trained AI Models on Copyrighted O’Reilly Media Books, Researchers Claim

Jane AustenBy Jane Austenabril 3, 2025No hay comentarios2 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
OpenAI Trained AI Models on Copyrighted O’Reilly Media Books, Researchers Claim
Share
Facebook Twitter LinkedIn Pinterest Email


OpenAI might have trained its artificial intelligence (AI) models on copyrighted content, according to a research paper. A recently published paper from the non-profit organisation AI Disclosures Project, the San Francisco-based AI firm’s recent large language models (LLMs) showed a higher recognition of copyrighted content compared to its older models. The researchers used a recently developed method called DE-COP to detect copyrighted content in the AI models’ training dataset. Notably, the study found that the GPT-4o mini was not trained on the specific copyrighted content.

Researchers Used DE-COP to Test OpenAI’s Training Dataset

The study, titled Beyond Public Access in LLM Pre-Training Data, was conducted to check if OpenAI’s AI models were trained on non-public book content. For the study, researchers focused on O’Reilly Media, a US online learning platform, which contains numerous copyrighted books. The founder of the platform, Tim O’Reilly, was also one of the co-authors of the study.

The researchers used DE-COP method to test whether the training data of the AI models contained copyrighted material. This is a relatively new test, introduced in a paper published in 2024. The method, also known as a membership inference attack, quizzes an AI model with a multiple-choice test to see whether it can identify copyrighted content from machine-generated paraphrased alternatives.

The researchers used Claude 3.5 Sonnet to paraphrase the copyrighted material. As many as 3,962 paragraph excerpts from 34 O’Reilly Media books were used for the test.

Based on the tests conducted, the researchers claimed to have found that the GPT-4o AI model showed the highest recognition of the copyrighted and paywalled O’Reilly book content with an 82 percent Area Under the Receiver Operating Characteristic Curve (AURUC) score. Notably, the AURUC score is part of the DE-COP method and is derived from the guess rates from the multiple-choice test.

The study also found that older OpenAI AI models, such as GPT-3.5 Turbo, showed lesser content recognition compared to GPT-4o, but still high enough to be significant. However, GPT-4o mini was found not to be trained on the paywalled O’Reilly Media books. The paper states the reason could be that the test is not effective against smaller language models.



Source link

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Jane Austen
  • Website

Related Posts

Bitcoin Core v30 allenta OP_RETURN: Alcuni miner S19 affrontano una pressione di booster

noviembre 17, 2025

Bitcoin Core v30: No es una actualización — es presionar el 「booster de eliminación de mineros」

noviembre 17, 2025

Pika Labs Launches Social AI Video App on iOS, Unveils New Audio-Driven Video Generation AI Model

agosto 12, 2025
Add A Comment
Leave A Reply Cancel Reply

Editors Picks

Fast fashion pioneer Forever 21 files for bankruptcy — again

marzo 18, 2025

Dow gains 350 points as stocks climb for 2nd day after S&P 500 enters correction

marzo 18, 2025

Yellow Creditors Have Own Plan to Share Trucker’s $550 Million

marzo 18, 2025

Alphabet in Talks to Buy Startup Wiz for $30 Billion, WSJ Says

marzo 18, 2025
Top Reviews
DD Noticias: Tu fuente de inspiración diaria
Facebook X (Twitter) Instagram Pinterest Vimeo YouTube
  • Home
  • Contact us
  • DMCA
  • Política de Privacidad
  • Publicidad en DD Noticias
  • Sobre Nosotros
  • Términos y Condiciones
© 2026 ddnoticias. Designed by ddnoticias.

Type above and press Enter to search. Press Esc to cancel.