The model
Technical report
The recipe and the measurement set-up, with the choices behind them.
The recipe
How was Walnoot 8B built?
Underneath Walnoot 8B sits Apertus-8B-2509, the open European base model from the Swiss AI Initiative. We built on it in three steps: continued pretraining, supervised fine-tuning and rejection fine-tuning.
Continued pretraining on Dutchcontinued pretraining
We trained the base model further up to eight billion tokens. The best result came at about four billion, after 3,816 steps of 1,048,576 tokens. This model rests on that point, followed by a cooldown of 64 steps. Seventy per cent of that text is Dutch, twenty per cent English and ten per cent code. The English and the code also come from the GPT-NL Public Corpus. They are in there as replay, because a model that only ever sees Dutch forgets what it could already do.
The sources for each item are on the data provenance page.
Learning to carry out instructionssupervised fine-tuning
A base model that has had continued pretraining completes text; it does not yet carry out an instruction. In this step it is shown 51,875 examples of an instruction with the desired answer, spread over twenty task categories. It sees that set twice (two epochs). The loss is computed on the answer tokens only. The instruction data is largely synthetic, with a scientific branch drawn from public annotation datasets. We compiled it ourselves and checked it on a sample.
The source table further down lists this as our own source.
Post-training with feedbackrejection fine-tuning
In the last step the model was trained further on answers it had produced itself. Earlier versions of the model answered questions from the same collection as the instruction data. That produced generation pools of 24,000, 12,000 and 15,400 answers. In the three task families with a checkable answer (classification, extraction and anonymisation) a rule-based check decided. It does not measure whether anything was invented. Outside those families two third-party models assessed each answer. If they disagreed, a third model decided by majority. At most one answer per question was kept. An approved answer longer than 1.8 times the median length in characters of the approved answers was still dropped.
That left 1,157 examples. On those the model was trained for one more epoch with supervised fine-tuning, at half the learning rate of step 02. The aim is that it more often delivers the form and the completeness the task calls for. There is no reward model and no reference model. We did try training on pairs of a good and a less good answer. That route is not in the model you download.
How that judging step worked is set out in the public summary of the training content.
The training run took less than a week and the first model was running within about a month. The configuration of each step is on GitHub, in the folder recept. For step 01 those are plan files. The values that step actually ran with are in the readme of that folder.
Base model
What does Walnoot 8B build on?
Walnoot 8B builds on Apertus. The full technical structure, the training steps and our own contributions are described in the model card.
Published
Which parts of this build are published?
- Training recipe
- Apache 2.0, public
- Evaluation configuration
- Apache 2.0, public, pinned to 17.6.0
- Evaluation logs
- Apache 2.0, public, raw jsonl
The corpus
Which sources are in the training data?
| Source | Basis | Processing |
|---|---|---|
| GPT-NL Public Corpus | CC-BY 4.0, publicly released without registration | Deduplicated against the rest of our mix. The mix is 70 per cent Dutch, 20 per cent English and 10 per cent code. The English and the code come from this corpus as well. The English comes from the subsets cc_english-pd, cc_loc-pd-books and cc_openalex, the code from cc_github_open_source. They are in there as deliberate replay against knowledge loss. |
| Instruction data (compiled by WAINUT) | Largely synthetic, made with a teacher model under Apache 2.0. Verbatim corpus passages fall under CC BY 4.0. The scientific branch comes from thirteen public datasets, each under the licence of that dataset. | Licence traced per external source. A gate on personal data for the corpus passages prepared for document tasks in September. No filter on personal data over the data as a whole. |
The continued-pretraining corpus also contains WAINUT's own source for preserving the formatting of text. Measured in characters it makes up 2.0 per cent of the corpus. The instruction data has a scientific branch drawn from thirteen public annotation datasets; the licence of each dataset is given in the public summary of the training content.
We left newspaper archives out on quality grounds. Text from behind a paywall is not in it either.
Measurement design
How was the measurement done?
- Harness
- EuroEval 17.6.0
- Regime
- test split, 10 iterations, bf16, zero failed instances
- Model
- Walnoot 8B Instruct 1.0
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
Why EuroEval?
EuroEval is the framework that GPT-NL chose itself. The version is pinned to 17.6.0 and the command is published. The full table is online, including the four tasks where we can claim nothing. On WikiLingua the metric changed between the two measurements, so we never put that task head to head. Leaving it out would be the real cherry-picking.
Reproduce
Show the command
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-splitDecision rule
When does a win count as a win?
A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.
GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.
Ceiling
Where does this model stop?
Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.
For broad knowledge and long reasoning a frontier model is the better tool.
Index