Replication package: What Citations Get Wrong — audit pipeline, anonymized citation-level judgments, and arbitration records
Description
Replication package for the paper What Citations Get Wrong: A Full-Corpus Audit of Reference Existence and Claim Support in a Major NLP Conference.
Contents. The two-layer audit pipeline (ingestion, GROBID parsing, existence verification against a local literature snapshot, two-stage support judgment, refutation-stance arbitration, reporting) with its dataset export and verification scripts; and the anonymized citation-level dataset: support judgments, arbitration records, and the L1 existence-triage residue, with all paper identifiers removed and no verbatim text, together with the aggregate reports the published numbers derive from. See README.md and dataset/README.md for a step-by-step reproduction of the headline figures.
Scope. The support-defect rate reported in the paper does not reproduce across re-runs, which is the paper's central finding; this archive reproduces every number computed from the frozen records, not the pipeline run itself.
Licensing. Dataset under CC BY 4.0; pipeline code under Apache-2.0 (see LICENSE files within the archive).
Note. This record is the object the numbers reported in the paper correspond to. The live repository at https://github.com/fim-ai/tuto may move ahead of it.
Files
Files
(499.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:d2c9ad09c43bc01e0e62f42dad013876
|
499.6 kB | Download |