Published May 12, 2026 | Version 1.0.0

Code: NIH DMPs LLM Evaluation Paper

  • 1. FAIR Data Innovation Hub
  • 2. California Medical Innovations Institute
  • 3. FAIR Data Innovations Hub
  • 4. California Digital Library
  • 5. University of California Office of the President

Description

About

Large language models are increasingly used to draft NIH Data Management Plans (DMPs), but their quality and policy alignment require careful evaluation. This repository contains the code for our systematic assessment of Llama 3.3 and GPT-4.1 using both automated metrics and human expert review. See the project inventory for related resources, including the paper and dataset.

Standards followed

The overall codebase is organized in alignment with the FAIR-BioRS guidelines. All Python code follows the PEP 8 conventions, including consistent formatting, inline comments, and docstrings. Project dependencies are fully captured in requirements.txt .

License

This work is licensed under the MIT License. See LICENSE  for more information.

Feedback and contribution

Use GitHub Issues to submit feedback, report problems, or suggest improvements.
You can also fork the repository and submit a Pull Request with your changes.

Files

nih-dmp-llm-evaluation-paper-code-v1.0.0.zip

Files (995.1 kB)

Name Size Download all
md5:0cca45c17a892ca51a384d1b4323a1c5
995.1 kB Preview Download

Additional details

Related works

Funding

U.S. National Science Foundation
Chan Zuckerberg Initiative (United States)

Dates

Available
2026