There is a newer version of the record available.

Published October 29, 2025 | Version v0.3.1

RISE-UNIBAS/humanities_data_benchmark

Description

This repository contains benchmark datasets (images), prompts, ground truths, and evaluation scripts for assessing the performance of large language models (LLMs) on humanities-related tasks. The suite is designed as a resource for researchers and practitioners interested in systematically evaluating how well various LLMs perform on digital humanities (DH) tasks involving visual materials. For detailed test results and model comparisons, visit our results dashboard at https://rise-services.rise.unibas.ch/benchmarks/. 

Notes

If you use this software, please cite it using the metadata from this file.

Files

RISE-UNIBAS/humanities_data_benchmark-v0.3.1.zip

Files (184.5 MB)

Name Size Download all
md5:b4017ae88e2ac24b89f645a1401115f4
184.5 MB Preview Download

Additional details

Related works