NASICON RAG Benchmark: Cross-LLM and Cross-Judge Validation Dataset and Code (v1.0)
Authors/Creators
Description
Dataset and code for the paper "Cross-LLM and Cross-Judge Validation of Retrieval-Augmented Generation for Materials Science: A Benchmark on NASICON-Type Solid Electrolytes" (submitted to Digital Discovery).
Contents: the NASICON v6 structured database (1,982 records; verbatim excerpts of non-CC-BY sources are withheld and replaced by DOI pointers, see database/LICENSE-DATA.md), the 52-question evaluation suite, all 44 evaluation-result JSONs (12 generators × multiple judges, with per-call timestamps and software versions), audit trails (output-budget incident; claim-decomposition judge-prompt variant), verbatim generator/judge prompts, source code, and analysis scripts that reproduce every table in the paper offline (analysis/reaggregate_rerun.py; see REPRODUCIBILITY.md for the three-tier reproducibility guide).
Code is released under the MIT License; data under CC BY 4.0. Files are restricted during peer review and will be made openly available upon acceptance.