Accountability
GPAI transparency
GPAI stands for general-purpose AI. This is where you find the three documents WAINUT publishes about Walnoot 8B and the address for your questions.
GPAI documents
What does WAINUT publish voluntarily?
The European AI Act asks anyone who provides a general-purpose AI model for three documents that can be public: a summary of the training content, a copyright policy and documentation for whoever builds the model into something else.
- Public summary of the training contentWhich data went into our own training step, per source and per kind, in the template of the European Commission.
- Copyright policyHow we deal with copyright and with reservations of rights under text and data mining, with the point of contact for rightsholders.
- Model documentation for downstream providersThe fields of the Model Documentation Form that are meant for downstream providers. We give the full form on request to the AI Office and to the national competent authorities.
These documents have not yet been reviewed by a legal adviser.
By its own measurement WAINUT is not a provider of a general-purpose AI model within the meaning of Regulation (EU) 2024/1689. We publish the three documents anyway, because openness about the data is the point of this project. WAINUT has not signed the Code of Practice for general-purpose AI models; we do follow it as guidance.
The funder prescribes this acknowledgement of the compute word for word.
We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-984 access to MareNostrum5 ACC hosted by BSC, SpainThe provider
Where can you turn with a question?
- Legal name
- WAiNuT B.V.
- Trading name
- WAINUT
- Chamber of Commerce number
- 93560915
- Address
- Rivium Westlaan 48, 2909LD Capelle aan den IJssel, Netherlands
- Point of contact
- support@walnoot.ai
- Website
- walnoot.ai
- Model
- Walnoot 8B Instruct 1.0
- Issued
- 28 September 2026
Rightsholders, downstream providers and anyone with a question about this model reach the same address.
Background
What else is on record about the model?
Which model is documented here?
- Model
- Walnoot 8B Instruct
- Version
- 1.0
- Licence
- Apache 2.0
- Base model
- Apertus
- Base model provenance
- Swiss AI Initiative (ETH Zürich and EPFL)
- Evaluation harness
- EuroEval 17.6.0
- Measurement regime
- test split, 10 iterations, bf16, zero failed instances
- Weights
- WAINUT/walnoot-8b-instruct
- Source code
- github.com/WAINUTAI/walnoot-8b-instruct
Where does the training data come from?
- GPT-NL Public Corpus
- BasisCC-BY 4.0, publicly released without registration
- Instruction data (compiled by WAINUT)
- BasisLargely synthetic, made with a teacher model under Apache 2.0. Verbatim corpus passages fall under CC BY 4.0. The scientific branch comes from thirteen public datasets, each under the licence of that dataset.
Excluded
- No newspaper archives.
- No text from behind a paywall.
- No data whose provenance is not fixed per source.
Which artefact is open and under which licence?
| Artefact | Licence | Access |
|---|---|---|
| Weights (8B) | LicenceApache 2.0 | AccessPublic, without registration |
| Evaluation logs | LicenceApache 2.0 | AccessPublic, raw jsonl |
| Training recipe | LicenceApache 2.0 | AccessPublic |
| Evaluation configuration | LicenceApache 2.0 | AccessPublic, pinned to 17.6.0 |
| Our Dutch corpus | LicenceDocumented per source | AccessProvenance public, source data not published by us |
| Training data of the base model | LicenceTerms per source | AccessDocumented by the Swiss AI Initiative, not ours to publish |
How were the figures measured?
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.
Under which law did the training run?
No verdict
European law for the instruction training, the post-training and the measurements. Those ran on MareNostrum5 in Barcelona, the European public supercomputer, via EuroHPC. Where you run the model is then up to you.
What falls outside the intended use
Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.
- Open-ended reasoning and long, complex analysis
- Creative work without a brief
- Code generation
- Anything where a mistake has immediate serious consequences without a person in between
The dossier
Which document is where?
- Model cardWhat Walnoot 8B is, what it was built for and where the limits lie.
- Data provenanceWhere the training data comes from, by source and by legal basis.
- Openness matrixWhat is open and under which licence, together with what is not ours.
- ScoreboardEvery EuroEval task next to the published figures from GPT-NL.
- Reproduce the resultsCheck every score yourself, with a pinned evaluation version.
- Evaluation logsThe raw jsonl behind every figure on the scoreboard.
- Technical reportThe recipe and the measurement set-up, with the choices behind them.
- LicenceApache 2.0, permanent and including commercial use.