Skip to main content

The evidence

Evaluation logs

We publish the raw measurement logs behind our figures on the scoreboard, so you can check them yourself.

Model
Walnoot 8B Instruct 1.0
Harness
EuroEval 17.6.0
Measurement regime
test split, 10 iterations, bf16, zero failed instances

Location

Where are the logs?

They are on GitHub, in the folder evaluatie. It holds the raw output of EuroEval 17.6.0 for all 10 tasks. You will find more there than in the table on the scoreboard, such as the 29.13 the model scores on conll-nl including MISC.

Recipe and logs on GitHub

Which tasks do they cover?

  1. dbrdSentimentMCCdirect
  2. squad-nlReading comprehensionEMdirect
  3. conll-nlNamed entity recognitionmicro-F1 (no-MISC)direct
  4. scala-nlLinguistic acceptabilityMCCdirect
  5. mmlu-nlKnowledgeMCCdirect
  6. hellaswag-nlCommonsense reasoningMCCdirect
  7. wikilingua-nlSummarisationChrF3++metric change
  8. duidelijke-taalPlain languageMETEORno publication
  9. valeu-nlEuropean valuesEuropeanValuesno publication
  10. mbbq-nlStereotypingbias-corrected accuracyno publication

Updated: 17 September 2026

Openness

Can you use the logs?

  • Evaluation logs: Apache 2.0, public, raw jsonl
  • Evaluation configuration: Apache 2.0, public, pinned to 17.6.0