Skip to main content

The evidence

Reproduce the results

Check every score yourself, with a pinned evaluation version.

Evaluation
EuroEval 17.6.0
Model
Walnoot 8B Instruct 1.0
Figures updated
17 September 2026
A nutcracker breaking a walnut open along the seam, with a few fragments beside it.

In three steps

How to check our figures yourself.

  1. Download the weights

    Model-id
    WAINUT/walnoot-8b-instruct
    Weights hash
    sha256:db131554ffb1fb9e0885855008b8ad4078077d4168e4c931e1024c2f9c586494sha256 of model.safetensorsThis is the hash of the weights file of this checkpoint. The file on Hugging Face has been read back and has the same hash, so your download should give exactly this sum.

    Weights on Hugging Face

  2. Run the command

    pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-split
    Measurement regime
    test split, 10 iterations, bf16, zero failed instances
  3. Compare your result with ours

    Below each figure is the 95% interval of our measurement.

    Walnoot 8B on the Dutch tasks of EuroEval 17.6.0. The figures from GPT-NL stand next to them on the scoreboard.
    TaskMetricWalnoot 8B
    dbrdMCCSentimentMCC90.5689.76 – 91.35
    squad-nlEMReading comprehensionEM62.0861.42 – 62.73
    conll-nlmicro-F1 (no-MISC)Named entity recognitionmicro-F1 (no-MISC)46.5244.49 – 48.55
    scala-nlMCCLinguistic acceptabilityMCC34.4131.00 – 37.81
    mmlu-nlMCCKnowledgeMCC37.1636.19 – 38.13
    hellaswag-nlMCCCommonsense reasoningMCC22.4620.73 – 24.19
    wikilingua-nlChrF3++SummarisationChrF3++34.1832.74 – 35.62
    duidelijke-taalMETEORPlain languageMETEOR52.9551.11 – 54.78
    valeu-nlEuropeanValuesEuropean valuesEuropeanValues2.021.41 – 2.64
    mbbq-nlbias-corrected accuracyStereotypingbias-corrected accuracy27.7226.17 – 29.27

    The raw jsonl behind each line is with the evaluation logs. The figures from GPT-NL stand next to them on the scoreboard.

Measurement design

How we measure and compare.

  • How were these figures measured?

    We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.

  • When does a task count as won?

    A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.

  • Why can you not check GPT-NL yourself?

    GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.

  • Is EuroEval not cherry-picking?

    EuroEval is the framework that GPT-NL chose itself. The version is pinned to 17.6.0 and the command is published. The full table is online, including the four tasks where we can claim nothing. On WikiLingua the metric changed between the two measurements, so we never put that task head to head. Leaving it out would be the real cherry-picking.