Published March 25, 2022 | Version 1.0

Construction and Evaluation of a High-Quality Corpus for Legal Intelligence Using Semiautomated Approaches

  • 1. University of North Texas

Description

This dataset is an extended version of from the aboved paper. The dataset contains more than 500,000 legal arguments which were automatically labeled by combining the following algorithms with majority vote:

  1. Fine-Tuned Legal BERT on training data involving few Paraphrased test samples. Accuracy: 0.834933 
  2. Fine-Tuned Legal BERT with custom algorithm involving a corruption of few train samples. Accuracy: 0.76775 
  3. Fine-Tuned Legal BERT Accuracy: 0.76199 
  4. Fine-Tuned BERT Accuracy: 0.76007 
  5. GAN-BERT Accuracy: 0.7480
  6. Co Training Accuracy: 0.66410 
  7. EM-LightGBM Accuracy: 0.6564 
  8. Pseudo Labeling Accuracy: 0.6339 

There are six labels in the dataset: 1) fact, 2) issue, 3)rule/law/holding, 4) analysis, 5) conclusion/opinion/answer, and 6) others. For more details about the labels and the algorithms, please read our paper onIEEE Transactions on Reliability

Notes

Email: Haihua.chen@unt.edu

Files

unt-legal-argument-mining-corpus(labeled).csv

Files (129.3 MB)

Name Size Download all
md5:0bfc6835e193a725d8861bddac576f7d
129.3 MB Preview Download

Additional details

Related works

Is cited by
Journal article: 10.1109/TR.2022.3156126 (DOI)