Skip to main content

The model

Model card

The model has 8.05 billion parameters and is released under Apache 2.0. It builds on Apertus and is made for bounded tasks.

Model
Walnoot 8B Instruct 1.0
Base model
Apertus
Licence
Apache 2.0

Identity

What is Walnoot 8B?

Walnoot 8B is a Dutch language model from WAINUT. You download the model and run it in your own environment, including for applications where data may not leave that environment. Everything we measure and publish, you can check yourself.

How to run Walnoot yourself

A walnut shell seen from very close up, with the grain lit obliquely from the side.

Specification

Which model is this?

Publisher
WAINUT
Model
Walnoot 8B Instruct
Version
1.0
Size
8.05 billion parameters
Base model
Apertus, Swiss AI Initiative (ETH Zürich and EPFL)
Licence
Apache 2.0
Language
Dutch
Evaluation
EuroEval 17.6.0

Recipe

How was Walnoot 8B built?

  1. European base

    Apertus-8B-2509 from the Swiss AI Initiative, published under Apache 2.0.

    What WAINUT added

    Nothing. This is the work of its makers.

  2. Continued pretraining on Dutch

    WAINUT trained that model further on Dutch-language material, with a cooling-down phase at the end. The mix was 70 per cent Dutch, 20 per cent English and 10 per cent code.

    What WAINUT added

    A Dutch base of our own. We do not publish it separately.

  3. Instruction layer

    On top of that the model learned to follow Dutch instructions: answering questions about a document, extracting information from a text and classifying texts.

    What WAINUT added

    An instruction layer of our own, with a chat template and a system prompt.

  4. Post-training with feedback

    Finally the model was steered on its own answers that a judging step approved.

    What WAINUT added

    The last step of the recipe. The technical term is rejection fine-tuning.

Is Walnoot 8B just an Apertus fine-tune?

Walnoot 8B builds on Apertus, but more happened than just a fine-tune. First the base model was trained further on billions of tokens of Dutch text. Then came an instruction layer of our own in Dutch and a final training step in which the model learned further from its own answers that an assessment step had approved. All of those steps are in the model you download.

An important difference lies in the data and in the choices we made there. We followed the line GPT-NL took on sovereignty: the provenance of sources is known, rights can be traced and the data used can be checked. That decided which data we did and did not use. It therefore also shapes what the model can ultimately do.

We publish the provenance of the Dutch data per source, the training recipe and the raw evaluation logs, so you can check most of the route.

Read the technical reportTo the data provenance

Intended use

What is the model intended for?

  1. Answering questions about a document

    Ask a question about a supplied document and have the answer taken from that document, with the source text alongside it.

    Reading comprehension, task squad-nl: 62.08 against the published 51 of GPT-NL.

  2. Extracting information from a text

    Pick names, places and organisations out of a Dutch text so that you can carry them over into a form or a register.

    Named entity recognition, task conll-nl: 46.52 against the published 36 of GPT-NL.

  3. Classifying texts

    Label incoming post, tickets and documents and route them to the right place.

    We measure classification through sentiment, task dbrd: 90.56 against the published 90 of GPT-NL. Under our own rule that is a tie.

  4. Drafting a standard text

    Draft a letter or a memo to a fixed pattern, using the details you supply.

    We have no measured figure for this.

  • Use summaries as a first draft and always check them before you use them any further.
  • Rewriting into plain language falls outside the intended use of Walnoot 8B.

View the full table

Ceiling

What falls outside the intended use

Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.

Not for

  • Open-ended reasoning and long, complex analysis
  • Creative work without a brief
  • Code generation
  • Anything where a mistake has immediate serious consequences without a person in between

A frontier model is the better tool for that.

Known limitations

What to know before you start

Three things stand out on first use, above all with English input. What we do ourselves is keep the input in Dutch and read through any summary.

  1. Repetition loop

    The model can get stuck in a repetition, producing the same sentence or the same turn of phrase over and over.

  2. Answer language

    English questions sometimes get a Dutch answer, while the system prompt prescribes that the model follows the language of the user.

  3. The name of the base model

    On English identity questions the name of the base model sometimes comes up spontaneously, while the identity policy only names it after further questioning.

We know these three and they are on the list for the next version.

Measuring

How was Walnoot 8B measured?

The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.

We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.

Show the command
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-split

Run the benchmark yourself

Why Walnoot

The research question and the outcome.

Our research question was whether a Dutch model, further trained on a European base model, could match GPT-NL on all six benchmarks that GPT-NL published, under comparable requirements for the provenance, the rights and the verifiability of the training data. It worked, with five scores above it and one tie.

What is open and under which licence?
The licence and the access per artefact. Where nothing else is stated, Apache 2.0 applies.
ArtefactLicenceAccess
Weights (8B)Apache 2.0Public, without registration
Evaluation logsApache 2.0Public, raw jsonl
Training recipeApache 2.0Public
Evaluation configurationApache 2.0Public, pinned to 17.6.0
Our Dutch corpusDocumented per sourceProvenance public, source data not published by us
Training data of the base modelTerms per sourceDocumented by the Swiss AI Initiative, not ours to publish

To the openness matrixRead the licenceTo the sovereignty test