WAINUT
Press kit
Figures, photos and logo, with the claim exactly as we phrase it.
Description
What is Walnoot 8B?
Walnoot 8B is a Dutch language model from WAINUT. You download the model and run it in your own environment, including for applications where data may not leave that environment. Everything we measure and publish, you can check yourself.
The claim
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd.
- Model
- Walnoot 8B Instruct
- Version
- 1.0
- Publisher
- WAINUT
- Licence
- Apache 2.0
- Size
- 8.05 billion parameters
- Base model
- Apertus
- Base model provenance
- Swiss AI Initiative (ETH Zürich and EPFL)
- Figures
- 17 September 2026
- Weights
- WAINUT/walnoot-8b-instruct
- Source code
- github.com/WAINUTAI/walnoot-8b-instruct
WAINUT
Who is WAINUT?
WAINUT is the Dutch AI company behind Walnoot, founded by Seleman Arefi and Luciano Currie.
Free to use in coverage of Walnoot and WAINUT, with the photo credit.
Imagery
Logo and mark
Measurement
What exactly does the claim say?
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.
When does a difference count as a win?
A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.
What was it measured on?
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.
Limits
What falls outside the intended use
- Open-ended reasoning and long, complex analysis
- Creative work without a brief
- Code generation
- Anything where a mistake has immediate serious consequences without a person in between
Why these limits?
Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.
Index
Where is the rest of the dossier?
- Model cardWhat Walnoot 8B is, what it was built for and where the limits lie.
- ScoreboardEvery EuroEval task next to the published figures from GPT-NL.
- Data provenanceWhere the training data comes from, by source and by legal basis.
- Frequently asked questionsWhat Walnoot is and how it relates to GPT-NL.



