Skip to search boxSkip to navigationSkip to main content

AI ‘News’ Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian

  • Giovanni Puccetti
    ,
  • ,
  • Chiara Alzetta
    ,
  • Felice Dell'Orletta
    ,
  • Andrea Esuli
  • Istituto di Scienza e Tecnologia dell’Informazione “A. Faedo”
    ,
  • ,
  • Italian Natural Language Processing Lab
Research Output:
Journal Article or Conference Article in Journal
Conference article
Peer-review

Open access

Publication Information

Output type

Research Output:
Journal Article or Conference Article in Journal
Conference article
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 15312–15338

Journal (Volume, Issue Number)

Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Publication milestones

  • Published - 2024

Publication status

Published - 2024

Abstract

Large Language Models (LLMs) are increasingly used as ‘content farm’ models (CFMs), to generate synthetic text that could pass for real news articles. This is already happening even for languages that do not have high-quality monolingual LLMs. We show that fine-tuning Llama (v1), mostly trained on English, on as little as 40K Italian news articles, is sufficient for producing news-like texts that native speakers of Italian struggle to identify as synthetic.We investigate three LLMs and three methods of detecting synthetic texts (log-likelihood, DetectGPT, and supervised classification), finding that they all perform better than human raters, but they are all impractical in the real world (requiring either access to token likelihood information or a large dataset of CFM texts). We also explore the possibility of creating a proxy CFM: an LLM fine-tuned on a similar dataset to one used by the real ‘content farm’. We find that even a small amount of fine-tuning data suffices for creating a successful detector, but we need to know which base LLM is used, which is a major challenge.Our results suggest that there are currently no practical methods for detecting synthetic news-like texts ‘in the wild’, while generating them is too easy. We highlight the urgency of more NLP research on this problem.

Access to documents

Related Event

Title

Conference on Association for Computational Linguistics

Event type

Conference

Degree of recognition

International event

Date

11/08/2024 - 16/08/2024

Location

BangkokThailand