Profile
Back to NewsBack
GitHub Trending 2 min
Reader Mode
EuroEval/EuroEval: The robust European language model benchmark.

EuroEval/EuroEval: The robust European language model benchmark.

17 hours ago

The robust European language model benchmark

(formerly known as ScandEval)


Documentation</a> PyPI Status</a> First paper</a> Second paper</a> License</a> LastCommit</a> Code Coverage</a> Contributor Covenant</a>

Maintainer

)

Installation and usage

See the documentation for more information.

Reproducing the evaluation datasets

All datasets used in this project are generated using the scripts located in the src/scripts/dataset_creation folder. To reproduce a dataset, run the corresponding script with the following command

uv run src/scripts/dataset_creation/<name-of-script>.py

Replace with the specific script you wish to execute, e.g.,

uv run src/scripts/dataset_creation/create_allocine.py

Contributors :pray:

A huge thank you to all the contributors who have helped make this project a success!

Contributor avatar for peter-sk Contributor avatar for AJDERS Contributor avatar for oliverkinch Contributor avatar for versae Contributor avatar for KennethEnevoldsen Contributor avatar for viggo-gascou Contributor avatar for mathiasesn Contributor avatar for Alkarex Contributor avatar for marksverdhei Contributor avatar for Mikeriess Contributor avatar for ThomasKluiters Contributor avatar for BramVanroy Contributor avatar for peregilk Contributor avatar for Rijgersberg Contributor avatar for duarteocarmo Contributor avatar for slowwavesleep Contributor avatar for mrkowalski Contributor avatar for sofiehb Contributor avatar for simonevanbruggen Contributor avatar for tvosch Contributor avatar for Touzen Contributor avatar for caldaibis Contributor avatar for SwekeR-463 Contributor avatar for N-essuno Contributor avatar for harderj Contributor avatar for lardinator Contributor avatar for lswiers Contributor avatar for avalyset Contributor avatar for pariidanDKE Contributor avatar for Biorrith Contributor avatar for FrejaThoresen Contributor avatar for rlrs Contributor avatar for jaideeppyne Contributor avatar for milos-plavsic Contributor avatar for Mr-Neutr0n Contributor avatar for djstrong

Contribute to EuroEval

We welcome contributions to EuroEval! Whether you're fixing bugs, adding features, or contributing new datasets, your help makes this project better for everyone.

for information on how to get started. to run or operate a community evaluation worker.
  • Adding datasets: If you're interested in adding a new dataset to EuroEval, we have
a dedicated guide with step-by-step instructions.

Special thanks

  • Thanks to Google for sponsoring Gemini credits as part of their
Google Cloud for Researchers Program.
  • Thanks @Mikeriess for evaluating many of the larger
models on the leaderboards.
  • Thanks to OpenAI for sponsoring OpenAI credits as part of their
Researcher Access Program.
  • Thanks to UWV and
KU Leuven for sponsoring the Azure OpenAI credits used to evaluate GPT-4-turbo in Dutch.
  • Thanks to Miðeind for sponsoring the OpenAI credits used to
evaluate GPT-4-turbo in Icelandic and Faroese.
  • Thanks to CHC for sponsoring the OpenAI credits used to evaluate
GPT-4-turbo in German.

Citing EuroEval

If you want to cite the framework then feel free to use this:

@inproceedings{smart2025encoder,
  title={Encoder vs decoder: Comparative analysis of encoder and decoder language models on multilingual NLU tasks},
  author={Smart, Dan Saattrup and Enevoldsen, Kenneth and Schneider-Kamp, Peter},
  booktitle={Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)},
  pages={561--572},
  year={2025}
}

@inproceedings{smart2023scandeval, author = {Smart, Dan Saattrup}, booktitle = {Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)}, month = may, pages = {185--201}, title = {{ScandEval: A Benchmark for Scandinavian Natural Language Processing}}, year = {2023} }

Chat with me