The evidence
Reproduce the results
Check every score yourself, with a pinned evaluation version.

In three steps
How to check our figures yourself.
Download the weights
- Model-id
- WAINUT/walnoot-8b-instruct
- Weights hash
- sha256:db131554ffb1fb9e0885855008b8ad4078077d4168e4c931e1024c2f9c586494sha256 of model.safetensorsThis is the hash of the weights file of this checkpoint. The file on Hugging Face has been read back and has the same hash, so your download should give exactly this sum.
Run the command
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-split- Measurement regime
- test split, 10 iterations, bf16, zero failed instances
Compare your result with ours
Below each figure is the 95% interval of our measurement.
Walnoot 8B on the Dutch tasks of EuroEval 17.6.0. The figures from GPT-NL stand next to them on the scoreboard. Task Metric Walnoot 8B dbrdMCCSentiment MCC 90.5689.76 – 91.35 squad-nlEMReading comprehension EM 62.0861.42 – 62.73 conll-nlmicro-F1 (no-MISC)Named entity recognition micro-F1 (no-MISC) 46.5244.49 – 48.55 scala-nlMCCLinguistic acceptability MCC 34.4131.00 – 37.81 mmlu-nlMCCKnowledge MCC 37.1636.19 – 38.13 hellaswag-nlMCCCommonsense reasoning MCC 22.4620.73 – 24.19 wikilingua-nlChrF3++Summarisation ChrF3++ 34.1832.74 – 35.62 duidelijke-taalMETEORPlain language METEOR 52.9551.11 – 54.78 valeu-nlEuropeanValuesEuropean values EuropeanValues 2.021.41 – 2.64 mbbq-nlbias-corrected accuracyStereotyping bias-corrected accuracy 27.7226.17 – 29.27 The raw jsonl behind each line is with the evaluation logs. The figures from GPT-NL stand next to them on the scoreboard.
Measurement design
How we measure and compare.
How were these figures measured?
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
When does a task count as won?
A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.
Why can you not check GPT-NL yourself?
GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.
Is EuroEval not cherry-picking?
EuroEval is the framework that GPT-NL chose itself. The version is pinned to 17.6.0 and the command is published. The full table is online, including the four tasks where we can claim nothing. On WikiLingua the metric changed between the two measurements, so we never put that task head to head. Leaving it out would be the real cherry-picking.