Published December 30, 2022 | Version 1.0

Annotated Pakistan Supreme Court Civil Appeal Judgments Dataset for Legal Named Entity Recognition

  • 1. Deep Learning Laboratory, National Center of Artificial Intelligence, Islamabad, Pakistan
  • 2. School of Electrical Engineering and Computer Sciences, National University of Science and Technology, Islamabad, Pakistan

Description

Overview

This dataset was prepared to support legal named entity recognition and automated information extraction from court judgments in Pakistan. The primary dataset contains token level annotations from English Civil Appeal judgments issued by the Supreme Court of Pakistan.

Court judgments are usually long and unstructured documents. Important information such as the names of people, organisations, courts, case numbers, dates, legal references, and monetary amounts may appear in different sections and formats. The purpose of this dataset is to make that information available in a structured form that can be used for developing and evaluating natural language processing systems.

Source of the data

The judgments used for the primary dataset were obtained from publicly available records of the Supreme Court of Pakistan. Civil Appeal judgments were selected because this category had the highest number of available cases on the Supreme Court website.

The downloaded records were processed through a validation and filtering workflow. Checksum validation was used to identify distinct records. Documents written in Urdu were removed from the collection. After preprocessing, the source paper reports a total of 214 distinct English Civil Appeal judgments. Individual judgments ranged from approximately 3 to 40 pages.

Annotation process

The annotation categories were selected after consultation with legal specialists. The judgments were annotated using Doccano, an open source text annotation platform. The resulting annotations were converted into the IOB sequence labelling format.

In the IOB format, the B prefix identifies the beginning of an entity, the I prefix identifies a token inside the same entity, and O identifies a token that does not belong to a labelled entity.

Each nonempty line in the Supreme Court dataset contains a token and its corresponding IOB label. Empty lines separate text sequences. The data is divided into training, validation, and test files to support model development and evaluation.

Named entity categories

The Supreme Court dataset contains fourteen legal entity categories:

Person

Names of judges, lawyers, litigants, officials, and other individuals mentioned in a judgment.

Location

Names of cities, areas, addresses, and other geographical locations.

Organisation

Names of public institutions, private organisations, businesses, universities, departments, authorities, and similar bodies.

Case number

The assigned number of the case being decided in the current judgment.

Respondent

The name of the person, group, organisation, or institution identified as a respondent in the current case.

Date

Complete dates appearing anywhere in the judgment.

Referred court

The name of a court that issued a judgment referred to in the current document.

Referred case

The title or citation of an earlier judgment referenced in the current case.

Legal reference

References to legislation, statutory provisions, rules, orders, codes, and other legal authorities.

Appeal court

The court that issued the judgment or order against which the current appeal was filed.

Appeal case number

The case number associated with the judgment or order being appealed.

Money

Monetary amounts, fines, compensation, debts, costs, and other financial values.

FIR number

A First Information Report number mentioned in the judgment.

Approved

Information indicating whether a judgment was approved for reporting.

Contents of the uploaded archive

The uploaded archive contains the following folders:

SCP

The Supreme Court of Pakistan legal NER dataset. This is the main dataset created and evaluated in the associated paper. It contains training, validation, and test files.

LHC

A Lahore High Court dataset used for comparative experiments. The associated paper explains that this dataset originated from previous research and was not created as the primary contribution of the current study.

CoNLL-2003

A general news domain named entity recognition benchmark. It is included for model development or comparison purposes and is not a Pakistani legal dataset.

Intended uses

The dataset can be used for research and development in:

  • Legal named entity recognition
  • Legal information extraction
  • Natural language processing of court judgments
  • Transformer model training and evaluation
  • Legal document indexing
  • Semantic legal search
  • Similar case retrieval
  • Legal question answering
  • Judicial knowledge base construction
  • Automated extraction of case metadata
  • Legal document anonymisation
  • Analysis of references between cases and courts

The extracted entities may also support downstream systems that organise judgments, connect related cases, populate structured legal databases, or remove personal information from public documents.

Benchmark models and results

The associated study evaluated three pretrained transformer models:

  • BERTBASE-uncased
  • BERTBASE-cased
  • LegalBERT

On the Supreme Court dataset, BERTBASE-uncased achieved an average F1 score of 92.47 percent. BERTBASE-cased achieved the highest average F1 score of 94.72 percent. LegalBERT achieved an average F1 score of 92.51 percent.

The results indicate that retaining letter casing was particularly useful for recognising entities in Pakistani court judgments.

Limitations

The primary dataset covers English language Civil Appeal judgments from the Supreme Court of Pakistan. It should not be treated as a complete representation of all Pakistani court proceedings.

The dataset may not directly generalise to criminal cases, constitutional petitions, family cases, tax matters, or judgments from other courts. Additional case categories may contain specialised entities that are not represented in the current annotation scheme.

Some entity categories have contextual overlap. A court name may be labelled as a referred court in one sentence and as an appeal court in another. A respondent may be an individual, a group, an organisation, or a government institution. Names of organisations and locations may also be similar. These factors can make individual labels difficult to distinguish.

The source judgments are public court records and may contain names or other personal information. Users should handle the material responsibly and follow the legal, ethical, and institutional requirements that apply to their research. The paper also recommends expanding the dataset to other categories of Supreme Court judgments in future work.

Citation

Users of this dataset should cite the associated conference paper:

Nida Ahmed, Seemab Latif, Rabia Irfan, Adnan Ul-Hasan, and Faisal Shafait. Comparison of Transformer Models for Information Extraction from Court Room Records in Pakistan. 2022 International Conference on Electrical, Computer, Communications and Mechatronics Engineering, ICECCME 2022. IEEE. DOI: 10.1109/ICECCME55909.2022.9988642.

Files

Information_Extraction_Pakistan_CourtRoom.zip

Files (3.3 MB)

Name Size Download all
md5:8c88feb5101bb5f47357415e42875ed0
3.3 MB Preview Download

Additional details

Dates

Issued
2022-12-30
Dataset version 1.0 released.