PCST: A Systematic Study of Extreme Low-Bit LLaMA-7B Compression Without Retraining
Description
This preprint presents a systematic experimental study of extreme
low-bit compression of LLaMA-7B without retraining, distillation,
pruning, or architectural changes.
More than 60 documented compression methods and variants were evaluated,
including Product Quantization, selective low-rank residuals,
activation- and Hessian-aware approaches, function-preserving gauge
transformations, additive quantization, sparse corrections, and
network-level output calibration.
The final PCST v5 artifact occupies 2.05 GiB. On a document-reset
WikiText-2 evaluation with 6,688 prediction positions, it achieves
perplexity 15.56 and 59.59% Top-10 overlap with the Q8 reference.
Although the result remains behind standard Q3_K_M quantization in
quality and speed, the study identifies reproducible positive results,
documents negative results, and demonstrates that local weight or block
reconstruction error often fails to predict end-to-end language-model
behavior.
Files
pcst_llama7b_compression_preprint_v1.pdf
Files
(157.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:91d72ccd988c68a4520e68dfddac5384
|
157.5 kB | Preview Download |