The model
Model card
The model has 8.05 billion parameters and is released under Apache 2.0. It builds on Apertus and is made for bounded tasks.
Identity
What is Walnoot 8B?
Walnoot 8B is a Dutch language model from WAINUT. You download the model and run it in your own environment, including for applications where data may not leave that environment. Everything we measure and publish, you can check yourself.

Specification
Which model is this?
- Publisher
- WAINUT
- Model
- Walnoot 8B Instruct
- Version
- 1.0
- Size
- 8.05 billion parameters
- Base model
- Apertus, Swiss AI Initiative (ETH Zürich and EPFL)
- Licence
- Apache 2.0
- Language
- Dutch
- Evaluation
- EuroEval 17.6.0
- Weights
- WAINUT/walnoot-8b-instruct
- Source code
- github.com/WAINUTAI/walnoot-8b-instruct
Recipe
How was Walnoot 8B built?
European base
Apertus-8B-2509 from the Swiss AI Initiative, published under Apache 2.0.
What WAINUT added
Nothing. This is the work of its makers.
Continued pretraining on Dutch
WAINUT trained that model further on Dutch-language material, with a cooling-down phase at the end. The mix was 70 per cent Dutch, 20 per cent English and 10 per cent code.
What WAINUT added
A Dutch base of our own. We do not publish it separately.
Instruction layer
On top of that the model learned to follow Dutch instructions: answering questions about a document, extracting information from a text and classifying texts.
What WAINUT added
An instruction layer of our own, with a chat template and a system prompt.
Post-training with feedback
Finally the model was steered on its own answers that a judging step approved.
What WAINUT added
The last step of the recipe. The technical term is rejection fine-tuning.
Is Walnoot 8B just an Apertus fine-tune?
Walnoot 8B builds on Apertus, but more happened than just a fine-tune. First the base model was trained further on billions of tokens of Dutch text. Then came an instruction layer of our own in Dutch and a final training step in which the model learned further from its own answers that an assessment step had approved. All of those steps are in the model you download.
An important difference lies in the data and in the choices we made there. We followed the line GPT-NL took on sovereignty: the provenance of sources is known, rights can be traced and the data used can be checked. That decided which data we did and did not use. It therefore also shapes what the model can ultimately do.
We publish the provenance of the Dutch data per source, the training recipe and the raw evaluation logs, so you can check most of the route.
Intended use
What is the model intended for?
Answering questions about a document
Ask a question about a supplied document and have the answer taken from that document, with the source text alongside it.
Reading comprehension, task squad-nl: 62.08 against the published 51 of GPT-NL.
Extracting information from a text
Pick names, places and organisations out of a Dutch text so that you can carry them over into a form or a register.
Named entity recognition, task conll-nl: 46.52 against the published 36 of GPT-NL.
Classifying texts
Label incoming post, tickets and documents and route them to the right place.
We measure classification through sentiment, task dbrd: 90.56 against the published 90 of GPT-NL. Under our own rule that is a tie.
Drafting a standard text
Draft a letter or a memo to a fixed pattern, using the details you supply.
We have no measured figure for this.
- Use summaries as a first draft and always check them before you use them any further.
- Rewriting into plain language falls outside the intended use of Walnoot 8B.
Ceiling
What falls outside the intended use
Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.
Not for
- Open-ended reasoning and long, complex analysis
- Creative work without a brief
- Code generation
- Anything where a mistake has immediate serious consequences without a person in between
A frontier model is the better tool for that.
Known limitations
What to know before you start
Three things stand out on first use, above all with English input. What we do ourselves is keep the input in Dutch and read through any summary.
Repetition loop
The model can get stuck in a repetition, producing the same sentence or the same turn of phrase over and over.
Answer language
English questions sometimes get a Dutch answer, while the system prompt prescribes that the model follows the language of the user.
The name of the base model
On English identity questions the name of the base model sometimes comes up spontaneously, while the identity policy only names it after further questioning.
We know these three and they are on the list for the next version.
Measuring
How was Walnoot 8B measured?
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
Show the command
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-splitWhy Walnoot
The research question and the outcome.
Our research question was whether a Dutch model, further trained on a European base model, could match GPT-NL on all six benchmarks that GPT-NL published, under comparable requirements for the provenance, the rights and the verifiability of the training data. It worked, with five scores above it and one tie.
What is open and under which licence?
| Artefact | Licence | Access |
|---|---|---|
| Weights (8B) | Apache 2.0 | Public, without registration |
| Evaluation logs | Apache 2.0 | Public, raw jsonl |
| Training recipe | Apache 2.0 | Public |
| Evaluation configuration | Apache 2.0 | Public, pinned to 17.6.0 |
| Our Dutch corpus | Documented per source | Provenance public, source data not published by us |
| Training data of the base model | Terms per source | Documented by the Swiss AI Initiative, not ours to publish |
To the openness matrixRead the licenceTo the sovereignty test