Does the optimal token misalignment threshold for maximizing alignment safety differ between Baichuan 2 and Vi
Description
Large Language Models (LLMs) have drawn a lot of attention due to their strong performance on a wide range of natural language tasks, since the release of ChatGPT in November 2022. LLMs' ability of general-purpose language understanding and generation is acquired by training billions of model's parameters on massive amounts of text data, as predicted by scaling laws cite\kaplan2020scaling,hoffmann2022training\. The research area of LLMs, while very recent, is evolving rapidly in many different ways. In this paper, we review some of the most prominent LLMs, including three popular LLM families
Research goal: Does the optimal token misalignment threshold for maximizing alignment safety differ between Baichuan 2 and Vicuna-13B when evaluated on harmful content generation datasets?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 9.0/10.
Notes
Files
paper.pdf
Files
(76.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:1960d8c49cb8c1041e4886aaf05c32e6
|
76.8 kB | Preview Download |