Skip to main navigation Skip to search Skip to main content

Crossing Domains without Labels: Distant Supervision for Term Extraction

Research output: Conference Article in Proceeding or Book/Report chapterArticle in proceedingsResearchpeer-review

Abstract

Abstract
Automatic Term Extraction (ATE) is a critical component in downstream NLP tasks such as document tagging, ontology construction and patent analysis. Current state-of-the-art methods require expensive human annotation and struggle with domain transfer, limiting their practical deployment. This highlights the need for more robust, scalable solutions and realistic evaluation settings. To address this, we introduce a comprehensive benchmark spanning seven diverse domains, enabling performance evaluation at both the document- and corpus-levels. Furthermore, we propose a robust LLM-based model that outperforms both supervised cross-domain encoder models and few-shot learning baselines and performs competitively with its GPT-4o teacher on this benchmark.The first step of our approach is generating psuedo-labels with this black-box LLM on general and scientific domains to ensure generalizability. Building on this data, we fine-tune the first LLMs for ATE. To further enhance document-level consistency, oftentimes needed for downstream tasks, we introduce lightweight post-hoc heuristics. Our approach exceeds previous approaches on 5/7 domains with an average improvement of 10 percentage points. We release our dataset and fine-tuned models to support future research in this area.
Original languageEnglish
Title of host publicationProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
EditorsSaloni Potdar, Lina Rojas-Barahona, Sebastien Montella
Number of pages13
Place of PublicationSuzhou (China)
PublisherAssociation for Computational Linguistics
Publication date1 Nov 2025
Pages1366-1378
ISBN (Print)979-8-89176-333-3
Publication statusPublished - 1 Nov 2025
EventEmpirical Methods in Natural Language Processing: Industry Track - Suzhou , China
Duration: 4 Nov 20259 Nov 2025
https://aclanthology.org/2025.emnlp-industry.0/

Conference

ConferenceEmpirical Methods in Natural Language Processing: Industry Track
Country/TerritoryChina
CitySuzhou
Period04/11/202509/11/2025
Internet address

Keywords

  • Automatic Term Extraction
  • Cross-domain Benchmarking in NLP
  • Large Language Models for ATE
  • Pseudo-labeling and Fine-tuning
  • Document-Level Consistency and Post-hoc Heuristics

Fingerprint

Dive into the research topics of 'Crossing Domains without Labels: Distant Supervision for Term Extraction'. Together they form a unique fingerprint.

Cite this