Kolibri-1 im Praxistest: NIS2, Kanzlei, IT-Dienstleister
Authors/Creators
Description
Ich teste das offene MoE-Sprachmodell Aleph Alpha Kolibri-1 (78,1 Mrd. Parameter, 3,46 Mrd. aktiv, FP8) auf einer einzelnen H200 unter vLLM 0.29.0 für NIS2/BSIG, Kanzlei mit Berufsgeheimnis und IT-Dienstleister. Grundlage sind 1.851 gewertete Läufe in drei Runden mit regelbasierter Auswertung. Das Modell ist schnell (rund 145 Token/s im Einzelstrom, erstes Token nach 0,4 s, 3.709 Token/s bei 64 parallelen kurzen Anfragen ohne Fehler) und widersteht Prompt-Injection in 39 von 40 Läufen. Bei deutschem Fachrecht ist es aus dem Kopf unzuverlässig: Die NIS2-Einstufung stimmt in 53 % der Läufe, alle drei Meldefristen in 20 von 50 Antworten, die einschlägige Vorschrift § 32 BSIG wurde nur einmal genannt. Die Fehler sind überwiegend Ausweichen, teils selbstsicheres Erfinden. Ohne Datum und Wissensstichtag im System-Prompt hält das Modell den eigenen Wissensstand für die Gegenwart; mit Datum, aber ohne Stichtag erfindet es Ereignisse. Mit dem Gesetzestext im Prompt liegt die Einstufung bei 96 %, die Triage bei 120 von 120 und die Fristen bei 50 von 50 Läufen. Praxisurteil: lokal betreibbarer, schneller Assistent für Entwürfe und IT-Aufgaben, für Rechtsfragen nur mit Gesetzestext im Prompt und menschlicher Prüfung. Stichprobe, kein Benchmark, ein Modell, keine Vergleichsmodelle.
Schlüsselwörter: Kolibri-1, NIS2, BSIG, Open-Weight-Modell, LLM-Evaluation, Halluzination, Prompt-Injection, vLLM
Abstract (English)
An independent practical test of the open-weight MoE model Aleph Alpha Kolibri-1 (78.1B total, 3.46B active, FP8) on a single H200 with vLLM 0.29.0, covering German NIS2/BSIG compliance, law-firm confidentiality and IT-service tasks (1,851 scored runs, rule-based scoring). The model is fast (about 145 tokens/s single stream, 3,709 tokens/s at 64 concurrent short requests, no errors) and resisted prompt injection in 39 of 40 runs. Without the statute text in the prompt it is unreliable on German regulatory law (NIS2 classification 53 % correct, all three reporting deadlines correct in 20 of 50 answers); with the text in the prompt it reaches 96 %, 120/120 and 50/50. Answer keys were produced with an AI assistant and are not legally reviewed. Small samples, one model, not a benchmark.
Abstract (Arabic)
Files
Beierle_Kolibri-1_Praxistest_Preprint.pdf
Files
(871.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:da3570b07bbfe6f568a71fb42c305146
|
871.7 kB | Preview Download |
Additional details
Software
- Repository URL
- https://github.com/sbeierle/kolibri1-praxistest