Skip to search boxSkip to navigationSkip to main content

From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification

  • Shanshan Xu
    ,
  • Santosh T.y.s.s
    ,
  • Oana Ichim
    ,
  • Isabella Risini
    ,
  • ,
  • Matthias Grabmaier
  • Technical University of Munich
    ,
  • Graduate Institute of International and Development Studies
    ,
  • Ruhr University Bochum
    ,
Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Pages from-to (Number of pages)

Pages 9558–9576

Publication milestones

  • Published - 12/2023

Publication status

Published - 12/2023

Publisher

Association for Computational Linguistics, United States

Publication IDs

  • Scopus: 85184812452

Host publication title

Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

Abstract

In legal NLP, Case Outcome Classification
(COC) must not only be accurate but also
trustworthy and explainable. Existing work
in explainable COC has been limited to an-
notations by a single expert. However, it is
well-known that lawyers may disagree in their
assessment of case facts. We hence collect
a novel dataset RAVE: Rationale Variation
in ECHR1, which is obtained from two ex-
perts in the domain of international human
rights law, for whom we observe weak agree-
ment. We study their disagreements and build a
two-level task-independent taxonomy, supple-
mented with COC-specific subcategories. We
quantitatively assess different taxonomy cate-
gories and find that disagreements mainly stem
from underspecification of the legal context,
which poses challenges given the typically lim-
ited granularity and noise in COC metadata. To
our knowledge, this is the first work in the legal
NLP that focuses on building a taxonomy over
human label variation. We further assess the ex-
plainablility of state-of-the-art COC models on
RAVE and observe limited agreement between
models and experts. Overall, our case study re-
veals hitherto underappreciated complexities in
creating benchmark datasets in legal NLP that
revolve around identifying aspects of a case’s
facts supposedly relevant to its outcome

Publication metrics

PlumX, opens in new tab

Citations
8
Captures
19

Related Event

Title

Conference on Empirical Methods in Natural Language Processing

Event type

Conference

Degree of recognition

International event

Date

06/12/2023 - 10/12/2023

Location

Resorts World Convention CentreSingapore