Skip to main content

Accountability

Frequently asked questions

What Walnoot is and how it relates to GPT-NL.

Model
Walnoot 8B Instruct 1.0
Harness
EuroEval 17.6.0
Updated
17 September 2026

10 basic questions

The basics first.

What is Walnoot 8B?

Walnoot 8B is an open-weight Dutch language model from WAINUT. You download it and run it on your own hardware or in your own cloud environment. The weights, the evaluation logs and the training recipes are under Apache 2.0. The model was continued pretrained and fine-tuned on top of Apertus, the open European base model that the Swiss AI Initiative at ETH Zürich and EPFL released under Apache 2.0.

Walnoot 8B is made for the work that makes up the largest part of AI use in an organisation: answering questions about a supplied document, classifying and routing incoming post, extracting data from a text and drafting a standard text. Reading comprehension, named entity recognition and classification carry a measured score in the published evaluation. Routing on the document types of an organisation itself is intended use and has not been measured separately. A summary from the model is a first draft that you read over. For rewriting into plain language we advise against the model. Both limits are in the model card as well.

Who is Walnoot 8B made for?

For Dutch organisations where restraint is an obligation. The first group is government: ministries, executive agencies, provinces, municipalities and public knowledge institutions. The second is regulated business, such as insurers, financial institutions, healthcare and energy. The third consists of closed environments, such as defence and critical infrastructure. What those groups share is a layered AI policy whose sensitive layer has no tooling today.

What is Walnoot 8B not for?

Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.

  • Open-ended reasoning and long, complex analysis
  • Creative work without a brief
  • Code generation
  • Anything where a mistake has immediate serious consequences without a person in between

A frontier model is the better tool for that. Most organisations push their entire workload through the most expensive and most exposed option there is, while a large part of that work needs nothing of the kind.

Read the model card

Can I run Walnoot 8B in my own environment?

Yes, that is what it is made for. The weights are under Apache 2.0 on Hugging Face. You run them on your own hardware or in your own cloud environment, on the ground you choose yourself. There is no mandatory call to an external API and no telemetry, so it also works fully disconnected. A model behind an API cannot do that. Continued pretraining on your own data can happen inside your own environment, without that data leaving it. What hardware you need and which system prompt belongs with the model are described on /proberen. How fast the model runs depends on your card and your memory.

What does this cost?

The model costs nothing. The weights, the logs and the recipes are under Apache 2.0, permanent and including commercial use, without registration and without a form. You pay for your own computing power and there is no meter running per token. WAINUT earns from the work around it: advice, staff training, implementation and staffing of AI talent. How that works per step is on Getting started.

Will this stay open?

What you have downloaded of the 8B stays yours under Apache 2.0, including for commercial use. The usage policy of the base model, the Apertus LLM Acceptable Use Policy, applies as well. Anyone may build on Walnoot, including a competitor. Openness is our philosophy. Later, larger models may come with different commercial terms. We make that distinction clear from the start.

Who has to arrange what if we deploy Walnoot 8B?

Hosting it yourself removes one objection. The assessment stays: the DPIA, the legal basis, the accuracy requirement and the human in the loop stay with the organisation that deploys the model. We supply the artefact that the assessment is about. An approval is not part of it.

What is WAINUT?

WAINUT is the Dutch AI company behind Walnoot, founded by two people. It builds open components that Dutch organisations can run in their own environment. Walnoot 8B is the model and NL-GOV-MCP fetches the Dutch sources the model works on. WAINUT put 50+ AI projects into production and trained 1000+ professionals on the work floor, at Dutch government bodies and regulated companies. Walnoot 8B is built on what those organisations need and therefore sized to their everyday work. The model is the free part. Deciding what you do with AI and getting it into production is the work.

Is there personal data in the training data?

We cannot rule that out. The training data as a whole was not filtered for personal data. Still, a sensitive part was removed and the data of the last training step was checked afterwards. It contained no Dutch citizen service numbers or bank account numbers. That check is coarse and does not show that the data is clean. The model can produce data that can be traced to a person. Whoever deploys it is the data controller for that. Do not use the model to remove personal data from documents either.

Is Walnoot 8B open source?

We call Walnoot 8B open-weight. The weights, the training recipe, the evaluation settings and the raw evaluation logs are under Apache 2.0, including for commercial use. We do not publish the training data itself. Open source would therefore promise more than there is. What is open per component is on /openheid.

The claim

What does Walnoot 8B claim?

The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.

When does a difference count as a win?

A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.

What was it measured on?

We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.

Why do you not measure GPT-NL yourselves?

GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.

View the scoreboardRecompute every score yourself

8 objections

The objections, with the answer

You only beat a low bar.

GPT-NL was the yardstick we fixed for ourselves beforehand. Our goal was clear: to perform at least as well on the comparable tasks. We did not move that bar during the build when results disappointed. On the scoreboard we show all ten tasks, including the four for which no direct comparison with GPT-NL is possible.

EuroEval is cherry-picking.

EuroEval is the framework that GPT-NL chose itself. The version is pinned to 17.6.0 and the command is published. The full table is online, including the four tasks where we can claim nothing. On WikiLingua the metric changed between the two measurements, so we never put that task head to head. Leaving it out would be the real cherry-picking.

That training time is marketing arithmetic.

The training run took less than a week. The first model was running within about a month. After that we refined it further. The release came almost three months after the first idea.

The 8B is just an Apertus fine-tune.

Walnoot 8B builds on Apertus, but more happened than just a fine-tune. First the base model was trained further on billions of tokens of Dutch text. Then came an instruction layer of our own in Dutch and a final training step in which the model learned further from its own answers that an assessment step had approved. All of those steps are in the model you download.

An important difference lies in the data and in the choices we made there. We followed the line GPT-NL took on sovereignty: the provenance of sources is known, rights can be traced and the data used can be checked. That decided which data we did and did not use. It therefore also shapes what the model can ultimately do.

We publish the provenance of the Dutch data per source, the training recipe and the raw evaluation logs, so you can check most of the route.

WAINUT has not trained a model from scratch. Part of the compute time for it was allocated only after the training run. Everything there was went into the 8B. That is the model people run.

Is a larger open model with a fine-tune on top not just as good?

On raw performance you are right. There are open-weight models that are larger than Walnoot 8B and that handle more on broad knowledge and long reasoning; that is stated here too, under the honest limits. A fine-tune helps on top of that, because the model will do your task better on your material.

What lies underneath is not touched by a fine-tune. It is a thin layer on top of a base model that has seen orders of magnitude more and what has been published about that does not change with it. How far you can look under the bonnet is determined by that base model.

That is what an assessor checks. A DPIA or a procurement review covers the whole chain and your fine-tune is the top layer of it. So look whether underneath the token count of a model there is a source list that you can check.

With Walnoot 8B the chain runs all the way down. Our own layer is listed per source on the provenance page, with the legal basis included. The base model is Apertus from the Swiss AI Initiative, with the documentation for it held by them. We build on a base model ourselves and that is why the choice of that foundation is the whole story.

Why Walnoot 8B and not a general open model that also runs locally?

General open-weight models can also be downloaded and run locally. Some are excellent and a number are considerably larger than Walnoot 8B. Running locally is a property of the whole category. Our answer has four parts.

Jurisdiction is the reason the category exists. If a good open model from another jurisdiction settled this question, no European or Dutch public money would go to European and Dutch models. Where a model comes from and who can change its terms are real questions for a public organisation. On top of that comes whether the chain can be explained to an assessor. Those questions remain, even if a general model speaks good Dutch.

Built and measured for Dutch work, sized to that work. Walnoot 8B was developed for everyday Dutch tasks and evaluated on them against a pinned evaluation version. The full results table is included, including where it falls short.

A package an assessor can assess. You get a named checkpoint in full precision, published evaluation logs with the command to check them, data provenance by source, an openness matrix per artefact and a model card. With a general model the buyer assembles that themselves, against a chain they have not documented and cannot pin.

A Dutch implementation path behind it. That path starts with the assessment beforehand and continues into domain tuning inside the organisation's own environment. After that come the training on the work floor and the anchoring via NL-GOV-MCP where Dutch sources matter. The model is the free part.

How does Walnoot compare to other open Dutch models?

There are more open Dutch models and that is good work by colleagues. What we do is publish what you need to check our work: the provenance per source, the recipe, the raw evaluation logs and all ten scores with their margin of error. The size is chosen for well-defined work. We do not put scores of other models next to ours, except the published figures of GPT-NL.

Where was Walnoot 8B trained?

The instruction training, the post-training and the measurements of Walnoot 8B ran on a European supercomputer, MareNostrum5 in Barcelona, via EuroHPC. The capacity is in the Netherlands as well, where Snellius at SURF in Amsterdam holds 352 H100s and the national consortium GPT-NL trained on 88 of them. That compute time is allocated through a call from NWO, for researchers with a permanent or tenure-track appointment at a Dutch knowledge institution. The 2026 call excludes private and public limited companies and cites the state aid rules itself in doing so. The European route is open to a company, because EuroHPC grants access to the AI Factories without requiring an appointment at a knowledge institution. Talent and plans are there in the Netherlands and so is the computing power. National access to it depends on an appointment at a knowledge institution. Anyone working from a company does not have one. The European route turned out to close that gap. Where you run the model is up to you.

Sovereignty

The sovereignty test

What does sovereign AI mean?

Across the AI sector, digital sovereignty comes up more and more often. For an organisation it means that the data stays inside and that it decides for itself where the model runs. With Walnoot 8B you can check four things yourself. You hold the artefact yourself, downloaded under Apache 2.0 and yours to keep. You choose the jurisdiction it runs in, on your own hardware or in your own cloud environment, on European soil or fully disconnected. Nothing has to leave your environment, because there is no mandatory call to an external API and we put no telemetry in it. And it is built in the Netherlands, on an open European base model. It is about this model; about the country we make no claim here. Where you run it is a different question from where it was built.

May you download the weights and keep them?

Passed

Yes. Apache 2.0, permanent, including commercial use. No registration, no revocable licence.

Do you know what it was trained on?

Passed

Yes, for our layer. Documented per source, with the legal basis and the processing steps included. The layer underneath is Apertus. That documentation belongs to the Swiss AI Initiative and sits with them.

Can you redo the training and the evaluation?

Passed

Yes, for the evaluation. The harness version is pinned and the command is on the site. You can trace the training through the recipe, but you cannot redo it exactly, because the instruction examples and the post-training data are not public.

Under which law did the training run?

No verdict

European law for the instruction training, the post-training and the measurements. Those ran on MareNostrum5 in Barcelona, the European public supercomputer, via EuroHPC. Where you run the model is then up to you.

View the sovereignty test