Curiosity-16: A 354.8M Parameter Large Language Model
Description
Despite their age, GPT-2 models remain among the most downloaded open-source large
language models. With a 2019 knowledge cutoff, and the tendency of GPT-2 models to hallucinate
or misinterpret, these models face significant drawbacks. Using GPT-2 Medium (354.8m
parameters) as the foundational model, we release Curiosity-16 (C16), a 354.8 million parameter
large language model that utilizes a two-phase supervised fine-tuning (SFT) pipeline for increased
domain-specific accuracy, reasoning, and recent knowledge injection. Using EleutherAI’s LM
Eval Harness, we evaluated Curiosity-16 and GPT-2 Medium on the HellaSwag and Massive
Multitask Language Understanding (MMLU) benchmarks. For zero-shot HellaSwag, Curiosity-16
shows a +0.29-percentage point increase in normalized accuracy over GPT-2 Medium, and for
MMLU, Curiosity-16 shows a +0.85-percentage point increase over GPT-2 Medium for
normalized accuracy. However, subject-level gains are more pronounced. Our targeted fine-tuning
pipeline affirms that model performance enhancements can be made with limited hardware and
publicly available resources.
Files
C16-Research-Paper-v1.pdf
Files
(338.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:357210eb8bff7b7d58b7ba131bdf30d0
|
338.3 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Model: https://github.com/ariankharazmi/Curiosity-16-LLM (URL)
- Model: https://huggingface.co/spaces/ariankharazmi/Curiosity-16 (URL)
- Model: https://huggingface.co/ariankharazmi/Curiosity-16 (URL)
Software
- Repository URL
- https://github.com/ariankharazmi/Curiosity-16-LLM
- Programming language
- Python
- Development Status
- Active