Annotated Pakistan Supreme Court Civil Appeal Judgments Dataset for Legal Named Entity Recognition
Authors/Creators
- 1. Deep Learning Laboratory, National Center of Artificial Intelligence, Islamabad, Pakistan
- 2. School of Electrical Engineering and Computer Sciences, National University of Science and Technology, Islamabad, Pakistan
Description
Overview
This dataset was prepared to support legal named entity recognition and automated information extraction from court judgments in Pakistan. The primary dataset contains token level annotations from English Civil Appeal judgments issued by the Supreme Court of Pakistan.
Court judgments are usually long and unstructured documents. Important information such as the names of people, organisations, courts, case numbers, dates, legal references, and monetary amounts may appear in different sections and formats. The purpose of this dataset is to make that information available in a structured form that can be used for developing and evaluating natural language processing systems.
Source of the data
The judgments used for the primary dataset were obtained from publicly available records of the Supreme Court of Pakistan. Civil Appeal judgments were selected because this category had the highest number of available cases on the Supreme Court website.
The downloaded records were processed through a validation and filtering workflow. Checksum validation was used to identify distinct records. Documents written in Urdu were removed from the collection. After preprocessing, the source paper reports a total of 214 distinct English Civil Appeal judgments. Individual judgments ranged from approximately 3 to 40 pages.
Annotation process
The annotation categories were selected after consultation with legal specialists. The judgments were annotated using Doccano, an open source text annotation platform. The resulting annotations were converted into the IOB sequence labelling format.
In the IOB format, the B prefix identifies the beginning of an entity, the I prefix identifies a token inside the same entity, and O identifies a token that does not belong to a labelled entity.
Each nonempty line in the Supreme Court dataset contains a token and its corresponding IOB label. Empty lines separate text sequences. The data is divided into training, validation, and test files to support model development and evaluation.
Named entity categories
The Supreme Court dataset contains fourteen legal entity categories:
Person
Names of judges, lawyers, litigants, officials, and other individuals mentioned in a judgment.
Location
Names of cities, areas, addresses, and other geographical locations.
Organisation
Names of public institutions, private organisations, businesses, universities, departments, authorities, and similar bodies.
Case number
The assigned number of the case being decided in the current judgment.
Respondent
The name of the person, group, organisation, or institution identified as a respondent in the current case.
Date
Complete dates appearing anywhere in the judgment.
Referred court
The name of a court that issued a judgment referred to in the current document.
Referred case
The title or citation of an earlier judgment referenced in the current case.
Legal reference
References to legislation, statutory provisions, rules, orders, codes, and other legal authorities.
Appeal court
The court that issued the judgment or order against which the current appeal was filed.
Appeal case number
The case number associated with the judgment or order being appealed.
Money
Monetary amounts, fines, compensation, debts, costs, and other financial values.
FIR number
A First Information Report number mentioned in the judgment.
Approved
Information indicating whether a judgment was approved for reporting.
Contents of the uploaded archive
The uploaded archive contains the following folders:
SCP
The Supreme Court of Pakistan legal NER dataset. This is the main dataset created and evaluated in the associated paper. It contains training, validation, and test files.
LHC
A Lahore High Court dataset used for comparative experiments. The associated paper explains that this dataset originated from previous research and was not created as the primary contribution of the current study.
CoNLL-2003
A general news domain named entity recognition benchmark. It is included for model development or comparison purposes and is not a Pakistani legal dataset.
Intended uses
The dataset can be used for research and development in:
- Legal named entity recognition
- Legal information extraction
- Natural language processing of court judgments
- Transformer model training and evaluation
- Legal document indexing
- Semantic legal search
- Similar case retrieval
- Legal question answering
- Judicial knowledge base construction
- Automated extraction of case metadata
- Legal document anonymisation
- Analysis of references between cases and courts
The extracted entities may also support downstream systems that organise judgments, connect related cases, populate structured legal databases, or remove personal information from public documents.
Benchmark models and results
The associated study evaluated three pretrained transformer models:
- BERTBASE-uncased
- BERTBASE-cased
- LegalBERT
On the Supreme Court dataset, BERTBASE-uncased achieved an average F1 score of 92.47 percent. BERTBASE-cased achieved the highest average F1 score of 94.72 percent. LegalBERT achieved an average F1 score of 92.51 percent.
The results indicate that retaining letter casing was particularly useful for recognising entities in Pakistani court judgments.
Limitations
The primary dataset covers English language Civil Appeal judgments from the Supreme Court of Pakistan. It should not be treated as a complete representation of all Pakistani court proceedings.
The dataset may not directly generalise to criminal cases, constitutional petitions, family cases, tax matters, or judgments from other courts. Additional case categories may contain specialised entities that are not represented in the current annotation scheme.
Some entity categories have contextual overlap. A court name may be labelled as a referred court in one sentence and as an appeal court in another. A respondent may be an individual, a group, an organisation, or a government institution. Names of organisations and locations may also be similar. These factors can make individual labels difficult to distinguish.
The source judgments are public court records and may contain names or other personal information. Users should handle the material responsibly and follow the legal, ethical, and institutional requirements that apply to their research. The paper also recommends expanding the dataset to other categories of Supreme Court judgments in future work.
Citation
Users of this dataset should cite the associated conference paper:
Nida Ahmed, Seemab Latif, Rabia Irfan, Adnan Ul-Hasan, and Faisal Shafait. Comparison of Transformer Models for Information Extraction from Court Room Records in Pakistan. 2022 International Conference on Electrical, Computer, Communications and Mechatronics Engineering, ICECCME 2022. IEEE. DOI: 10.1109/ICECCME55909.2022.9988642.
Files
Information_Extraction_Pakistan_CourtRoom.zip
Files
(3.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:8c88feb5101bb5f47357415e42875ed0
|
3.3 MB | Preview Download |
Additional details
Dates
- Issued
-
2022-12-30Dataset version 1.0 released.