KAI-95M: Benchmarking An Efficient Transformer Alternative
Authors/Creators
Contributors
Project manager:
Project member (2):
Description
We introduce KAI-95M, a highly capable 95M parameter language model that leverages a hybrid closed- form continuous time (CfC) and mLSTM model architecture that surpasses the performance of LLMs ∼16x its size (GPT-2 1.5B) at greatly reduced computational cost. KAI-95M has been designed to resolve the memory limitations of previous CfC architectures via the introduction of mLSTM memory units and a gating mechanism for memory unit regulation. This allows KAI-95M to meet the challenge of on-prem and on-device language model deployment without the need for GPUs to run inference based tasks. Saving money, saving compute, and providing users with a model that has demonstrated its ability to accomplish specific language modeling tasks with a parameter efficiency that is, on average, ∼22x greater than the models it is benchmarked against herein.
Designed to ensure seamless fine-tuning across a plethora of tasks, KAI-95M represents an exciting new chapter in the era of specialized small language models (SLMs). Please contact us for research validation and other technical inquiries.
Files
main.pdf
Files
(186.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:305fbf8c24f6f7fed71fd4371f83d662
|
186.9 kB | Preview Download |