GPAI documents
Model documentation for downstream providers
The fields of the Model Documentation Form that are meant for downstream providers.
Document dated 28 September 2026. Model version 1.0, released on 28 September 2026.
Published voluntarily by WAINUT; see the statement in this document.
This document has not yet been reviewed by a legal adviser.
Quotations from EU documents are reproduced verbatim in English.
This is the information for downstream providers from the Model Documentation Form for Walnoot 8B Instruct; the Form as a whole is handed over on request to the AI Office or to the national competent authority.
Filled in by WAINUT using the Model Documentation Form annexed to the Transparency Chapter of the Code of Practice for general-purpose AI models.
Status. On its own measurement, recorded in WAINUT's own internal assessment of its duties for this model, WAINUT is not a provider of a general-purpose AI model within the meaning of Regulation (EU) 2024/1689 (the AI Act), because the training compute of its modification stays far below the indicative threshold of the Commission Guidelines. Even for a provider, the three transparency measures do not apply to a model released under a free and open-source licence without systemic risk: "In accordance with Article 53(2) AI Act, these Measures do not apply to providers of general-purpose AI models released under a free and open-source license that satisfy the conditions specified in that provision, unless the model is a general-purpose AI model with systemic risk." WAINUT fills the Form in voluntarily, for the same reason the base model publisher did.
Conventions. No internal paths, host names, keys or storage addresses appear here. No prompt, record or answer text from the training data appears here. Figures are written with the thousands separator of English usage. Dutch quotations are reproduced verbatim; the English rendering in square brackets after each of them is WAINUT's own translation and is not authoritative.
Date this document was last updated:
28 September 2026
General information
Legal name for the model provider:
Provider: WAiNuT B.V. (registered legal name), trading as WAINUT, a private limited company registered with the Netherlands Chamber of Commerce (KVK) under number 93560915, Rivium Westlaan 48, 2909LD Capelle aan den IJssel, the Netherlands. Point of contact: support@walnoot.ai. Website: https://walnoot.ai. Establishment number: 000059092947.
The contact point above is the address for questions about this Form, for downstream providers, for rightsholders and for complaints.
Model name:
Walnoot 8B Instruct, version 1.0. Distributed as WAINUT/walnoot-8b-instruct. Publicly available version covered by this documentation: version 1.0, whose last training step is post-training with feedback (rejection fine-tuning, RFT), the name the model card gives that step. It is ordinary supervised fine-tuning on the model's own answers that an assessment step had approved; the published training recipe states in its own heading that there is no reference model and no preference pairs. No other version has been made publicly available.
Release date:
28 September 2026.
Union market release date:
28 September 2026, the same date as the release date above.
Model dependencies:
swiss-ai/Apertus-8B-2509, published by the Swiss AI Initiative under the Apache License 2.0. Walnoot 8B Instruct is the result of continued pretraining, instruction tuning and a post-training step applied to that model.
The publisher of the base model also applies a separate acceptable use policy with its own version number, the Apertus LLM Acceptable Use Policy version 1.0 of 1 September 2025, which can change independently of the licence. That name, version and date come from the model card of the base model itself, in its gated-access field, which carries the heading "Apertus LLM Acceptable Use Policy" followed by "(1.0 | September 1, 2025)". WAINUT's own internal licence assessment pins the same policy, but to the 70B teacher model rather than to the base model, so that assessment is not the source for this entry.
A second model of the same publisher, swiss-ai/Apertus-70B-Instruct-2509 under the Apache License 2.0, was used as a teacher to generate the synthetic answers in the instruction data. It generated the answers, not the instruction data as a whole: the source documents of part of that data come from the third-party public corpus; the scientific branch comes from existing public annotation datasets. It is not a dependency of the released weights.
Model properties
Model architecture:
A decoder-only transformer of the Apertus family. The configuration file shipped with the weights declares the architecture as ApertusForCausalLM with model type apertus. It has 32 transformer layers, a hidden width of 4,096 and an intermediate width of 21,504. Attention is grouped-query attention, with 32 attention heads over 8 key-value heads and with query and key normalisation. The activation is xIELU and the normalisation is RMS normalisation with an epsilon of 1e-05. Positional encoding is rotary of the llama3 type, with a base of 12,000,000 and a scaling factor of 8.0 over an original context of 8,192, which gives an architecture context of 65,536 positions. The vocabulary holds 131,072 tokens, the weights are bfloat16 and the input and output embeddings are not tied.
The model keeps the architecture of the base model swiss-ai/Apertus-8B-2509. The modification changed weights only. That is read from the configuration file shipped with the weights and is not derived from the training method. Total parameters 8.05 billion. Architecture context length 65,536, of which 32,768 is served.
Input modalities:
- [X] Text
- [ ] Images
- [ ] Audio
- [ ] Video
If any other please specify: none.
Maximum input size: text, 32,768 tokens as served. The architecture carries 65,536.
Output modalities:
- [X] Text
- [ ] Images
- [ ] Audio
- [ ] Video
If any other please specify: none.
Maximum output size: N/A. WAINUT sets no maximum output length. What bounds a generation is the served context length of 32,768 tokens.
Total model size (range):
- [ ] 1-500M
- [ ] 500M-5B
- [X] 5B-15B
- [ ] 15B-50B
- [ ] 50B-100B
- [ ] 100B-500B
- [ ] 500B-1T
- [ ] >1T
Methods of distribution and licenses
Distribution channels (for downstream providers):
The model is available to downstream providers as open weights at the model page https://huggingface.co/WAINUT/walnoot-8b-instruct, with weights-level access. There is no gated access, no application procedure and no API operated by WAINUT. The public repository on GitHub, which carries the training recipe, the evaluation settings and job script, the provenance documents and the decontamination documents, is https://github.com/WAINUTAI/walnoot-8b-instruct.
This column of the Form, the public summary of the training content and the copyright policy are published in English at https://walnoot.ai/en/transparantie and in Dutch at https://walnoot.ai/transparantie. The contact point for downstream providers, for rightsholders, for complaints and for questions is support@walnoot.ai.
Why WAINUT fills this Form in at all. On its own measurement, recorded in WAINUT's own internal assessment of its duties for this model, WAINUT is not a provider of a general-purpose AI model within the meaning of the AI Act. WAINUT publishes these documents voluntarily. WAINUT does not sign the Code of Practice and follows it as guidance.
License (for downstream providers):
A free and open-source licence, the Apache License 2.0, under which the model can be openly shared and downstream providers can freely access, use, modify and redistribute it or modified versions of it. WAINUT imposes no additional restriction on use, no non-commercial clause, no research-only clause and no user threshold.
Note for downstream providers: the base model publisher maintains a separate acceptable use policy with its own version number, which can change independently of the licence. WAINUT's own release carries the plain Apache License 2.0 without an addendum.
Additional assets made available:
| Asset | How it can be accessed | Licence |
|---|---|---|
| Model weights, safetensors and bf16 GGUF | model page, https://huggingface.co/WAINUT/walnoot-8b-instruct | Apache License 2.0 |
| Model card, with data provenance, evaluation and limitations | model page, as above | Apache License 2.0 |
| NOTICE with the attribution list per source | model page, as above | as above |
| Weights manifest with sizes and hashes | model page, as above | as above |
| System prompt file, 1,527 bytes, with a pinned sha256 | model page, as above | as above |
| Training recipe in three parts, evaluation settings and job script, provenance documents, decontamination documents | public repository on GitHub, https://github.com/WAINUTAI/walnoot-8b-instruct | Apache License 2.0 |
| Training data | not made available. The model card states: "Er wordt geen trainingsdata gepubliceerd. Wij publiceren de brondata niet; wat openligt is de herkomst per bron." [No training data is published. We do not publish the source data; what is public is the provenance per source.] | not applicable |
Use
Acceptable Use Policy:
WAINUT publishes no acceptable use policy of its own. The restrictions and warnings that exist are carried by the model card. Measure 1.4 of the Copyright Chapter allows a free and open-source release to alert users in the accompanying documentation instead of in an acceptable use policy. The model card carries a sentence that alerts users to the prohibition on infringing use.
Intended uses:
Walnoot 8B Instruct is a language model for Dutch. It is intended as an aid in work with Dutch business and legal texts. It is not a search engine, not a fact base and not a substitute for professional judgement. It was trained on Dutch instruction tasks: answering a question against a supplied document, extracting data from a text, classifying texts and drafting standard texts. Summarisation is included as well, but only as a first draft that the user reads back.
Limited and prohibited uses, taken from the model card:
- The model must not be used to remove personal data from documents.
- Rewriting into plain language does not work well, summarisation is weak. The model card records four numbered limitations in total.
- The model has not been tested for safety in the sense of a structured red-team round. The model card carries that point with its results rather than in the numbered list: the bias task in the public evaluation battery touches bias and is not a safety assessment.
- A user who deploys the model in an environment with personal data remains the data controller and cannot rely on WAINUT's measurements.
Type and nature of AI systems in which the general-purpose AI model can be integrated:
Text-in, text-out systems in the Dutch language: assistants and drafting aids for public-sector and administrative text, extraction of data from text and classification of texts, question answering over supplied documents. Summarisation fits only as a first draft that the user reads back. The model is not suited to rewriting into plain language. The model was trained with a fixed system prompt, which is part of the integration; the EuroEval measurements ran without a system prompt.
The model is not suitable as a component in a system whose purpose is to detect or remove personal data, because the model card forbids that use. The model has not been assessed for use in any high-risk setting under the AI Act. This Form makes no statement about such use.
Technical means for model integration:
Weights in safetensors format, plus a bf16 GGUF conversion for runtimes that read that format. The fixed system prompt ships as a file of 1,527 bytes with a pinned sha256; the model card states that it belongs with the model, because without it the model behaves differently from what it was tuned for. Neutral generation settings used for measurement: temperature 0.0, top_p 1.0, top_k 0, min_p 0.0, repeat_penalty 1.0.
Required hardware:
WAINUT sets no hardware requirement for inference. The anchor that does exist is the weights file of 16,106,727,320 bytes in bf16, so what is needed is an accelerator with memory above that size plus runtime overhead, or a runtime that quantises the weights while it loads them. Two distribution artefacts are published, the safetensors file and a bf16 GGUF conversion. WAINUT publishes no quantised variant.
Required software:
WAINUT sets no inference software requirement either. What can be given instead is the serving stack in which the published measurements were made, read from its own container image: vllm 0.24.0+cu129, transformers 5.14.1, torch 2.11.0+cu129, euroeval 17.6.0, xgrammar 0.2.3, flashinfer-python 0.6.12, python 3.12. That is a measured environment and not a requirement. The bf16 GGUF file runs in any runtime that reads that format.
For completeness, the training image, which is likewise not an inference requirement: axolotl 0.17.0.dev0, transformers 5.9.0, torch 2.10.0+cu128, trl 1.5.1, peft 0.19.1, deepspeed 0.18.9, flash-attn 2.8.1, python 3.12. That image carries a second python tree with other versions in it; the eight figures here come from the first tree.
Information on the data used for training, testing, and validation
Data type/modality:
- [X] Text
- [ ] Images
- [ ] Audio
- [ ] Video
Data provenance:
Select all that apply:
- [ ] Web crawling
- [ ] Private non-publicly available datasets obtained from third parties
- [ ] User data
- [X] Publicly available datasets
- [X] Data collected through other means
- [X] Synthetic data that is not publicly accessible (when created directly by or on behalf of the provider)
- If any other please specify: none.
Web crawling is not ticked because WAINUT ran no crawler of its own. Web-derived material reaches the model only through the third-party public corpus and through the base model, which is described in its publisher's own Summary.
Nobody crawled on WAINUT's behalf either. The material came in through the open dataset repository of the publisher of the corpus and through public datasets. That is a statement of WAINUT itself and it is the only source there is for it. On both archive domains behind the two subsets discussed below there is no robots.txt; that is a measured fact about those domains and it is not an argument, because WAINUT fetched nothing there. The reservation of rights that counts is therefore the one the publisher of that dataset had to observe.
Data curation methodologies:
- Length and language filters. Grounding documents had to sit between 400 and 6,000 characters with a language score of at least 0.80.
- Filtering and deduplication over the whole corpus: 806,473 documents dropped over nineteen reasons, together 9,137,134,433 characters, with near-duplicate removal the largest reason at 377,787 drops, language 141,535 drops.
- Exclusion of newspaper archives.
- Exclusion of eight parquet files in one subset that a virus scanner had flagged.
- Licence triage per source, traced to primary sources.
- Decontamination against the public evaluation battery: before training on an earlier build of the corpus and on WAINUT's own source for preserving the formatting of text in the corpus; after training on the instruction data and the post-training data that were actually trained. The build of the corpus on which the model was trained has not been measured again as a whole.
- Rejection of 16,776 candidate passages on personal data, among the verbatim corpus passages set aside for document tasks in the instruction data, with a re-run of the gate returning zero hits on 1,600 passages. The gate covered only those candidate passages; the rest of the training data did not go through it.
On the licence value of two corpus subsets. At the pinned revision on which the continued pretraining ran, 26 of the 28 subsets used carry an explicit licence value in the per-document metadata and two archive subsets carry the value "unknown". Together those two hold 81,169 documents, 117,179,121 characters and 37,011,724 estimated tokens: 0.43 percent of the estimated tokens and 0.41 percent of the characters across all 28 subsets. The value stands in WAINUT's own provenance manifest. The script that built the corpus wrote it there from the licence column per document in the parquet files of the corpus at that revision; the per-subset statistics files that the publisher ships in the dataset repository carry the same value at that revision. On 7 September 2026, after that revision, the publisher set both subsets to "public-domain". WAINUT has not been able to establish where the value came from; in the description of that commit the publisher writes that the value looks like a gap in the metadata rather than an intentional value. Three facts weigh against reading that value as a restriction. The publisher of the corpus records both subsets as public-domain in its own collection metadata. Both archives place their open data under CC0. And what was used is text from character recognition and not image material, so no separate right in a scan or a photograph comes into play. Both subsets stay in the training data, because WAINUT's copyright policy works at source level and the corpus stands there as one source under CC BY 4.0. The account with the sources quoted is in the public summary of the training content at https://walnoot.ai/en/transparantie.
Compute acknowledgement, required verbatim in every publication about this model
Required acknowledgement, verbatim:
"We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-984 access to MareNostrum5 ACC hosted by BSC, Spain"