Published February 13, 2026
| Version v0.4.11
Software
Open
EleutherAI/lm-evaluation-harness: v0.4.11
Authors/Creators
- Lintang Sutawika1
- Hailey Schoelkopf
- Leo Gao
- Baber Abbasi
- Stella Biderman2
- Jonathan Tow
- ben fattori
- Charles Lovering
- farzanehnakhaee70
- Jason Phang
- Anish Thite3
- Fazz
- Thomas Wang4
- Niklas
- Aflah5
- sdtblck
- nopperl
- gakada
- tttyuntian
- researcher2
- Julen Etxaniz6
- Chris7
- James A. Michaelov8
- Hanwool Albert Lee9
- Janna
- Leonid Sinev
- Khalid
- Kiersten Stokes10
- Zdeněk Kasner11
- KonradSzafer
- 1. Language Technologies Institute, CMU
- 2. Booz Allen Hamilton, EleutherAI
- 3. playscape.gg
- 4. MistralAI
- 5. Max Planck Institute for Software Systems: MPI SWS
- 6. Hitz Zentroa EHU
- 7. @azurro
- 8. MIT
- 9. Shinhan Securities Co.
- 10. Open Source Developer @ IBM
- 11. Charles University
Description
v0.4.11 Release Notes
Minor release. Stay tuned for bigger changes next release.
New Platform Support
- Windows ML Backend — Native Windows ML inference support by @chapsiru and @chemwolf6922 in #3470, #3564, #3565
New Benchmarks & Tasks
- BEAR knowledge probe by @plonerma in #3496
Task Version Changes
The following tasks have updated versions. Results from a previous task versions may not be directly comparable. See the linked PRs or individual task READMEs for changelogs.
afrobench_belebele (all variants): 2 → 3 in #3551
evalita_llm: 0.0 → 0.1 in #3551
include (all 90 language variants): 0.0 → 0.1 in #3551
mgsm_direct (all 11 language variants): 3.0 → 4.0 by @LakshyaChaudhry in #3574
Fixes & Improvements
- Fixed SQuAD v2 evaluation by @HydrogenSulfate in #3535
- Fixed MasakhaNEWS tasks — replaced non-existent
headline_textfield withheadlineby @Mr-Neutr0n in #3567 - Fixed incorrect task configs by @baberabb in #3552
- Replaced
eval()withast.literal_evalin task configs for safer parsing by @baberabb in #3577 - Fixed SGLang duplicate registration error by @enpimashin in #3543
- Restored
hf_transferimport check by @baberabb in #3563 - Fixed
modify_gen_kwargscall in vLLM VLMs by @hmellor in #3573 - Refactored vLLM
gen_kwargsnormalization inline tomodify_gen_kwargs; fixed cachedgen_kwargsmutation by @baberabb in #3582 - Fixed README for task-listing CLI command by @UltimateJupiter in #3545
- Updated dependencies by @baberabb in #3546
New Contributors
- @HydrogenSulfate made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3535
- @UltimateJupiter made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3545
- @enpimashin made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3543
- @chapsiru made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3470
- @chemwolf6922 made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3565
- @plonerma made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3496
- @hmellor made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3573
- @Mr-Neutr0n made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3567
- @LakshyaChaudhry made their first contribution in https://github.com/EleutherAI/lm-evaluation-harness/pull/3574
Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.10...v0.4.11
Files
EleutherAI/lm-evaluation-harness-v0.4.11.zip
Files
(10.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:23cdc2f2230f0c437647eb6b2e5944ed
|
10.6 MB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/EleutherAI/lm-evaluation-harness/tree/v0.4.11 (URL)
Software
- Repository URL
- https://github.com/EleutherAI/lm-evaluation-harness