GPAI documents
Public summary of the training content
Which data went into our own training step, per source and per kind.
Document dated 28 September 2026. Model version 1.0, released on 28 September 2026.
Published voluntarily by WAINUT; see the statement in this document.
This document has not yet been reviewed by a legal adviser.
Quotations from EU documents are reproduced verbatim in English.
Filled in by WAINUT using the Template annexed to the Commission's Explanatory Notice for the Public Summary of Training Content for general-purpose AI models required by Article 53(1)(d) of Regulation (EU) 2024/1689 (the AI Act).
Status. By its own assessment WAINUT is not a provider of a general-purpose AI model within the meaning of the AI Act, because the training compute of its modification stays far below the indicative threshold of the Commission Guidelines. That assessment is worked out in WAINUT's internal records. WAINUT nevertheless fills in this Template voluntarily, because openness about data provenance is the point of this project. WAINUT is not a signatory to the Code of Practice for general-purpose AI models. It follows that Code as a guideline without signing it.
Conventions. Field names below are reproduced verbatim from the Template. No internal paths, host names, keys or storage addresses appear in this document. No prompt, record or answer text from the training data appears in this document. Dutch quotations are reproduced verbatim; the English rendering in square brackets after each of them is WAINUT's own translation and is not authoritative.
Version of the Summary:
Version 1.0, published with the model on 28 September 2026. No earlier version of this Summary has been published, so no links to earlier versions are given.
Last update:
28/09/26
1. General information
1.1 Provider identification
Provider name and contact details:
Provider: WAiNuT B.V. (registered legal name), trading as WAINUT, a private limited company registered with the Netherlands Chamber of Commerce (KVK) under number 93560915, Rivium Westlaan 48, 2909LD Capelle aan den IJssel, the Netherlands. Point of contact: support@walnoot.ai. Website: https://walnoot.ai.
Where these come from: the registered legal name, the legal form, the Chamber of Commerce number and the address are the entry of the company in the Dutch trade register; the registered legal name is reproduced above in the spelling of that register. The e-mail address is the contact point of WAINUT for rightsholders, for downstream providers, for complaints and for questions. No telephone number and no named contact person are given here: the contact point above is the whole of it.
Authorised representative name and contact details:
N/A. WAINUT is established in the Netherlands, so the obligation in Article 54 AI Act does not apply. Point (48) of the Commission Guidelines states that "providers established or located within a third country must appoint an authorised representative established in the Union before placing a general-purpose AI model on the Union market (Article 54 AI Act)".
1.2 Model identification
Versioned model name(s):
Walnoot 8B Instruct, version 1.0, distributed as WAINUT/walnoot-8b-instruct at https://huggingface.co/WAINUT/walnoot-8b-instruct. The last training step of the released model is post-training with feedback (rejection fine-tuning, RFT), the name the model card gives that step. It is ordinary supervised fine-tuning on the model's own answers that an assessment step had approved. Licence: Apache License 2.0. Total parameters: 8.05 billion. Languages: Dutch and English. Served context length 32,768. The model card published with the weights is the authoritative public description of the model.
Model dependencies:
Walnoot 8B Instruct is the result of a modification, including fine-tuning, of swiss-ai/Apertus-8B-2509, an open European base model published by the Swiss AI Initiative under the Apache License 2.0. The base model publisher has published its own Public Summary of Training Content, which covers the data used to train the base model. That Summary, not this one, describes the base model's training data.
That Summary is published at https://huggingface.co/swiss-ai/Apertus-70B-2509/blob/main/Apertus_EU_Public_Summary.pdf. The model card of the base model links to it under the heading "EU AI Act Transparency Documentation and Code of Practice". It is marked version V1 with a last update of 1 September 2025 and it names Apertus-8B, Apertus-70B, Apertus-8B-Instruct and Apertus-70B-Instruct as the models it covers. The same publisher publishes a byte identical copy of that file in its own public legal repository.
A second model of the same publisher, swiss-ai/Apertus-70B-Instruct-2509, was used as a teacher to generate synthetic instruction data. That is reported under Section 2.5 of this Summary, not here, because Walnoot 8B Instruct is not a modification of that model.
Date of placement of the model on the Union market:
28/09/26. Walnoot 8B Instruct is issued on 28 September 2026 and is placed on the Union market on that same date. The model card published with the weights carries that date.
1.3 Modalities, overall training data size and other characteristics
The Template introduces this Section as follows: "This Section requires general information about the overall training data after pre-processing and before the training of the model." The figures below concern WAINUT's own modification only.
| Modality | Training data size | Types of content |
|---|---|---|
| [X] Text | [X] 1 billion to 10 trillions tokens. Alternatively, specify the approximate size in a different measurement unit: approximately 4.07 billion tokens of continued pretraining, 3,816 steps of the main phase plus 64 steps of the cooldown phase at 1,048,576 tokens per step, derived from the step count and not counted with the tokenizer, plus 40,957,043 measured tokens in 51,875 instruction records, plus 1,177,826 measured tokens in 1,157 records in the post-training layer. | Dutch public-sector and administrative texts, court decisions, archival material, encyclopaedic and general web text released under open licences, English text, source code, synthetic Dutch instruction data, synthetic chat pairs from a source of WAINUT's own for preserving the formatting of text (made with the same teacher model as the instruction data, about 2.0 percent of the characters of the continued pretraining), existing public scientific annotation datasets in English. |
| [ ] Image | not applicable | none |
| [ ] Audio | not applicable | none |
| [ ] Video | not applicable | none |
| [ ] Other | not applicable | none |
The token figure for the continued pretraining is derived from the number of optimiser steps. The released model rests on the saved state at step 3,816 of the main phase, which is 4,001,366,016 tokens, followed by a separate cooldown phase of 64 steps on the same corpus, which is 67,108,864 tokens. Together that is 4,068,474,880 tokens. One step is 1,048,576 tokens, the product of micro batch size 8, gradient accumulation 4, 8 ranks and sequence length 4,096. The figure is not a count with the tokenizer. The pretraining run went on to step 7,632, about 8 billion tokens, in line with the token budget of the corpus build. The best result lay at step 3,816 and the released model rests on that saved state. The size of the corpus from which those tokens were drawn is an estimate derived from character counts and not a measurement with the tokenizer. The corpus build used a token budget of 8,000,000,000 with an oversample factor of 1.10; the per-subset estimate over 28 subsets sums to 8,615,679,298 estimated tokens.
The two other figures are measurements with the tokenizer of the released model, made on 23 September 2026. The instruction data that was trained holds 51,875 records and 40,957,043 tokens. The post-training layer holds 1,157 records and 1,177,826 tokens. Both counts cover the text of every turn in every record, the system turn included. Counted as the two files render under the chat template of the released model, so that the role markers and the beginning of sequence token are counted as well, the figures are 41,299,228 tokens and 1,185,925 tokens. The repeated system turn accounts for 13,364,632 tokens of the first figure and for 488,254 tokens of the second; without it the counts are 27,592,411 tokens and 689,572 tokens.
Latest date of data acquisition/collection for model training:
09/2026. That is the latest acquisition of data for this modification. The scientific branch of the instruction data was read in one acquisition in 09/2026 and the last of WAINUT's own generation runs was executed in 09/2026. For the continued pretraining the third-party Dutch corpus was read earlier, in 07/2026.
That corpus was read at revision 1e17d21afa6518939523fbdcc1fbb3167b671725, with lastModified 2026-05-04T12:42:07Z.
There was no explicit revision pin in the code that streamed the instruction data. The continued pretraining ran on the pinned revision above. The repository received two more commits after that revision. The commit of 24 August 2026 changed only the README. The commit of 7 September 2026 changed the licence value of the two archive subsets of the Noord-Hollands Archief and Het Utrechts Archief, in their two statistics files and in five parquet files; according to the description of that commit only the licence column changed in those parquet files. The other 26 subsets that were used are byte-identical at the revision above and at the head of the repository as it stood on 27 September 2026. Adding an explicit pin stands in WAINUT's internal records as a recommendation.
Not recorded per source set for the fourteen public scientific annotation datasets. They did not enter one by one. The whole scientific branch was read in a single acquisition of a third-party instruction mixture at one pinned revision. Every record of that branch in the instruction data that was trained carries that pin as its recorded source: allenai/SciRIFF via ai2-adapt-dev/tulu_v3.9_sciriff_10k in swiss-ai/apertus-sft-mixture at revision 51998bcb2f80bc3ad1779bdfe5ebdc4d4df9f624. That acquisition took place in 09/2026. The pinned revision itself was last modified on 19 September 2025.
The model is not trained continuously on new or dynamic data after release.
Description of the linguistic characteristics of the overall training data:
Dutch is the target language of the modification. Dutch is an official language of the Union.
The continued pretraining corpus was built to a token budget mix of 70 percent Dutch, 20 percent English, 10 percent source code, with an oversample factor of 1.10 and characters per token of 3.166 for Dutch, 3.761 for English and 3.599 for code. A language filter was applied during the corpus build: 141,535 documents were dropped on language. Documents used to ground the instruction data had to reach a language score of at least 0.80, with a length window of 400 to 6,000 characters.
The instruction layer is Dutch. The scientific branch of the instruction data is English.
Other relevant characteristics of the overall training data:
- No personal data filter ran over the training data as a whole. The model card states this plainly: "Persoonsgegevens. De trainingsdata is als geheel niet door een PII-filter gegaan." [Personal data. The training data as a whole has not been through a personal data filter.] What does exist is an after-the-fact measurement on the post-training layer, reported under Section 3.2.
- Newspaper archives were excluded, on quality grounds rather than on rights grounds. The model card states: "Krantenarchieven zijn uitgesloten. De grond is kwaliteit en geen recht: het gaat om historisch materiaal met een verouderde spelling en veel ruis uit tekstherkenning." [Newspaper archives are excluded. The ground is quality and not rights: this is historical material with outdated spelling and much noise from character recognition.]
- Eight parquet files flagged by a virus scanner in one source subset were excluded through an explicit exclusion list, an own tightening above the source.
- Filtering and deduplication. Across the whole corpus 806,473 documents were dropped over nineteen reasons, together 9,137,134,433 characters, with near-duplicate removal the largest single reason at 377,787 drops.
- Decontamination against the public evaluation battery was run in three measurements. Two ran before training, on an earlier build of the corpus and on WAINUT's own source of synthetic chat pairs for preserving the formatting of text in the corpus. The third ran after training, on the instruction data and the post-training data that were actually trained. The build of the corpus on which the model was trained has not been measured again as a whole. See Section 3.3.
- No training data is published. The model card states: "Er wordt geen trainingsdata gepubliceerd. Wij publiceren de brondata niet; wat openligt is de herkomst per bron." [No training data is published. We do not publish the source data; what is public is the provenance per source.] What is published is the provenance per source.
- At the revision that was used, two of the 28 corpus subsets carry the licence value "unknown" in the metadata per document. Together those two are 81,169 documents, 117,179,121 characters and 37,011,724 estimated tokens: 0.43 percent of the estimated tokens and 0.41 percent of the characters over all 28 subsets. On 7 September 2026, after that revision, the publisher set both to "public-domain". What that value rests on, what the sources themselves say and what WAINUT concluded is set out in Section 2.1 under "Additional comments".
Additional comments (optional):
The snapshot of the third-party source repository holds 8,320 files across 32 subsets and 948.14 GB. WAINUT's internal records of that snapshot carry no count of the subsets used. How many were used is reported in Section 2.1 below, on the authority of the model card.
Before training the corpus went through a decontamination check of its own. That check is not about the evaluation test sets. It is about a retention set that WAINUT keeps apart to measure whether the model loses knowledge during the continued pretraining. The check passed with one amendment. One gate in that check returned a failed outcome; on the retention set the check touched 2,754 of 6,244 records, 44.11 percent. The model card publishes that outcome. That outcome has not been withdrawn.
2. List of data sources
2.1 Publicly available datasets
Have you used publicly available datasets to train the model?
- [X] Yes
- [ ] No
Modalities covered:
- [X] Text
- [ ] Image
- [ ] Audio
- [ ] Video
List of large publicly available datasets:
GPT-NL Public Corpus, repository identifier GPT-NL/GPT-NL_Public_Corpus, released by the GPT-NL consortium under CC BY 4.0. Pinned revision 1e17d21afa6518939523fbdcc1fbb3167b671725 (lastModified 2026-05-04T12:42:07Z). The corpus is public and ungated.
This is the only dataset that passes the Template's threshold for "large", which the Template defines as follows: "A dataset is considered to be large if the total data size for any one of the modalities contained in the dataset exceeds 3% of the size of all publicly available datasets for that modality used for training."
What was taken from it and how it was selected: 28 of the 32 subsets were used. After filtering the built corpus holds 116 shards, 3,448,292 training documents, 29,240,230,710 characters, plus 5,451 validation documents. Newspaper archives were excluded. Eight virus-scanner-flagged parquet files in one subset were excluded. The subsets cover Dutch government and administrative publications, court decisions, archival collections, encyclopaedic and general web text released under open licences, English text and open-source code. The largest Dutch-language subset in the provenance manifest holds 576,454 documents and 8,563,896,056 characters under CC BY.
The repository is at https://huggingface.co/datasets/GPT-NL/GPT-NL_Public_Corpus.
Beside the corpus, the continued pretraining holds a small source of WAINUT's own, about 2.0 percent of the characters. It consists of synthetic chat pairs, made with the same teacher model as the instruction data, so that the model keeps the formatting of text. That source holds 311,105 records with 585,781,778 characters, 9.0 percent of the documents of the built corpus. The figures of the built corpus above include it.
General description of other publicly available datasets not listed above:
Fourteen public scientific annotation datasets were drawn on to build the scientific branch of the instruction data. Thirteen of them are in the data that was trained; the fourteenth, AnatEM, is explained under point 2 below. Per the four points the Template asks for:
- Modalities: text only.
- Nature of the content: biomedical and scientific annotation data, including question answering, entity annotation, summarisation targets and clinical trial eligibility text. The licence of each was established from a primary source. They are, with their licence: BioASQ (CC BY 2.5), MedMentions (CC0 1.0), Chia (CC BY 4.0), SciTLDR (Apache 2.0), QASA (MIT), COVID-QA (Apache 2.0), JNLPBA (CC BY 3.0), NCBI Disease (public domain), AnatEM (CC BY-SA 3.0), Qasper (CC BY 4.0), PubMedQA (MIT), MatSci Text Corpus (MIT), NLM-Chem (CC0 1.0), LINNAEUS (BSD). That is fourteen source sets, from which sixteen tasks were drawn. AnatEM is the one share-alike source set in that branch. WAINUT's own check of that branch records it as delivering 168 records and establishes its licence as CC BY-SA 3.0 rather than the CC BY the paper's table claims. Those 168 records do not reach the instruction data of 51,875 records that was trained. The branch was delivered as 2,659 records over nine licence rows, one of which is that CC BY-SA 3.0 row with 168 records. Inside the instruction data that was trained the branch holds 2,491 records, which is those 2,659 records without the CC BY-SA 3.0 row. The eight other rows match record for record: CC BY 2.5 812, Apache 2.0 516, MIT 412, CC0 1.0 281, CC BY 4.0 279, public domain 171, CC BY 3.0 16 and BSD 4, together 2,491. A count over the licence field of that instruction data, made on 23 September 2026, returns ten rows and no share-alike row among them. No share-alike obligation therefore rests on this release. The attribution list published with the weights names thirteen annotation sets and correctly does not name AnatEM.
- Linguistic characteristics: English.
- Approximate start and end date of collection: Not recorded per source set. The branch as a whole was read in one acquisition in 09/2026, as stated under "Latest date of data acquisition/collection for model training" above.
Additional comments (optional):
Licence selection rule. Per external source the licence was traced to a primary source, meaning the licence file or the dataset card of the source set itself, or the appendix of that set's own paper. The tenth row of the triage is WAINUT's own identity layer, which comes from no external source at all. The model card states: "Per externe bron is de licentie herleid uit een primaire bron, dus uit het licentiebestand of de datasetkaart van die bron zelf." [Per external source the licence has been traced to a primary source, that is to the licence file or the dataset card of that source itself.] Of 9,722 candidate records 4,421 passed that licence gate. 2,659 remained after all the later quality gates. The tightening of the rule moved the figure from 4,702 to 4,421 at the licence gate, dropping three source sets with 281 records between them. One of those three is the case that prompted the rule: no primary source permits commercial use, while the most specific source available, the intermediate resource from which the corpus itself draws that set, states a non-commercial licence where the paper's own table claimed CC BY. At the licence gate as a whole more fell away than those three: five named source sets, plus the sixteen source sets for which the paper's table fixes no licence at all so that no primary source was sought.
Those sixteen hold 4,749 records between them: CRAFT-Chem 305, SciCite 305, HealthVer 296, Scientific Papers 564, ACL ARC 293, LitCOVID 284, ChemSum 283, Drug Combinations 283, Lay Summarisation 542, ChemProt 531, EBM-NLP PICO 251, CHEMDNER 245, CovidFact 234, BioCreative V CDR 143, GNormPlus 96 and SciERC 94.
Licence composition of the instruction mix that was actually trained. The mix under the released model holds 51,875 records. The source-level licence triage returns ten rows: 49,134 records under Apache-2.0-clean, 812 under CC BY 2.5, 516 under Apache 2.0, 412 under MIT, 281 under CC0 1.0, 279 under CC BY 4.0, 250 proprietary to WAINUT, 171 public domain, 16 under CC BY 3.0 and 4 under BSD. Apache-2.0-clean is not a licence of an external source but a label WAINUT derives for the generated material of the mix. The records that rest on a literal passage from the corpus carry that label as well; the passage itself falls under CC BY 4.0, the licence of the corpus. The corpus is attributed in the NOTICE file. No non-commercial licence, no share-alike licence, no GPL and no research-only licence occurs in the mix.
Two corpus subsets with the licence value "unknown", disclosed here in good faith. At the revision that was used, two of the 28 corpus subsets carry that value in the metadata per document: a subset of 7,995 documents from the Noord-Hollands Archief and a subset of 73,174 documents from Het Utrechts Archief, two Dutch regional archives. Together they are 81,169 documents, 117,179,121 characters and 37,011,724 estimated tokens: 0.43 percent of the estimated tokens and 0.41 percent of the characters over all 28 subsets. The other 26 subsets each carry an explicit licence value. WAINUT's own copyright policy requires that an "unknown" value stops the pipeline for investigation. That investigation was carried out on 22 September 2026. Its outcome follows.
Where the value stands and where WAINUT took it from. The value stands in WAINUT's own provenance manifest, the file herkomst/provenance-manifest.json that goes to the public repository on GitHub at https://github.com/WAINUTAI/walnoot-8b-instruct in the folder herkomst, on exactly two of the 28 subsets. WAINUT did not assign it: the script that built the corpus wrote it into the manifest from the licence column per document in the parquet files of the corpus, where at the revision that was used the value reads "unknown" for these two subsets. The per-subset statistics files that the publisher of the corpus ships in the dataset repository carry the same value at that revision for these two subsets and "Public domain" for the neighbouring archive subsets. On 7 September 2026 at 11:34 UTC, after that revision, the publisher set both subsets to "public-domain" in a new commit, revision 1b4408c2e78fced22626123c1e8ad840d5183f12. That commit changed the two statistics files and, according to its description, the licence column of every document in both subsets. The dataset card of the corpus does not carry "unknown" among the values it documents for that field; the values it names are "public domain/CC-0/CC-by".
What the publisher of the corpus says about these two subsets. In the collection metadata of the corpus, a file of 452,087 bytes with sha256 46fa6d006f80cd482d8a950bb1c26c55d86d48463832457a29b4cb7842ffc88e as retrieved on 22 September 2026, both archives carry verbatim "license": {"type": ["public-domain"]}. The row counts in that file are equal to WAINUT's own snapshot, so these are demonstrably the same two subsets. That file can change, which is why its size, its checksum and the date of retrieval are given here. The publisher describes the Utrecht part as collections older than 100 years, with a temporal coverage before 1950 of 1.0. The Noord-Holland part it describes as council minutes of one municipality, digitised around 2015 by optical character recognition.
What the two archives themselves say. Both archives publish their open data under CC0. The page with the open data of the Noord-Hollands Archief, fetched fresh on 22 September 2026 with no retrieval date recorded per page, states verbatim: "Personen en organisaties kunnen de open data van het Noord-Hollands Archief vrijelijk gebruiken, al dan niet voor commerciële doeleinden. Er geldt de licentie CC0 (no rights reserved), dat wil zeggen iedereen mag de gegevens vrij gebruiken zonder toestemming te vragen. Bronvermelding is niet nodig, maar wordt wel op prijs gesteld." [Persons and organisations may freely use the open data of the Noord-Hollands Archief, whether or not for commercial purposes. The CC0 licence applies (no rights reserved), that is to say anyone may use the data freely without asking permission. Attribution is not required but is appreciated.] The page with the open data of Het Utrechts Archief, fetched fresh in the same way and likewise with no retrieval date recorded per page, states verbatim: "Personen en organisaties kunnen de open data vrijelijk gebruiken, al dan niet voor commerciële doeleinden. Wel zijn er enkele spelregels. Voor de nu vrijgegeven gegevens geldt de licentie CC0." [Persons and organisations may freely use the open data, whether or not for commercial purposes. There are, however, a few ground rules. For the data released so far the CC0 licence applies.] One qualification belongs with those two quotations: they are made about the open data of the archives, that is about their archival finding aids, not in so many words about the text of the documents themselves.
What WAINUT could not establish. Why the value reads "unknown" at the revision that was used is explained nowhere at that revision. In the description of its commit of 7 September 2026 the publisher writes that the value looks like a gap in the metadata rather than an intentional value. Whether the publisher took the value over from the archive or entered it itself for want of a field is stated neither on the dataset card nor in the summary of the corpus paper. Three facts weigh against reading that value as a restriction on re-use. The publisher counts both collections as public domain in its own collection metadata. Both archives publish their open data under CC0 and nowhere forbid re-use of this material. And what was used is text from character recognition and not image material, so no separate right in a scan or a photograph comes into play: the Noord-Holland part is council minutes of one municipality, digitised around 2015 by character recognition; the Utrecht part is on the publisher's own figure entirely material from before 1950. WAINUT has not walked through the two subsets document by document at the source.
The outcome. Both subsets stay in the training data. The copyright policy of WAINUT works at the level of the source; at that level the corpus is one source under CC BY 4.0: the publisher states that licence in the repository tag, on the dataset card and in the corpus paper. The publisher counts both subsets as public domain in its own collection metadata. Removing the two parts is in any case no longer possible, because the model has been trained and which documents were drawn was not recorded, so no list of the documents used exists. WAINUT reports that here rather than leave it to be found.
2.2 Private non-publicly available datasets obtained from third parties
2.2.1 Datasets commercially licensed by rightsholders or their representatives
Have you used datasets commercially licensed by rightsholders or their representatives to train the model?
- [ ] Yes
- [X] No
Modalities covered: none.
2.2.2 Private datasets obtained from other third parties
Have you used private datasets obtained from other third parties to train the model?
- [ ] Yes
- [X] No
Modalities covered: none.
If publicly known, list private datasets obtained from other third parties:
N/A.
General description of non-publicly known private datasets obtained from third parties:
N/A.
Additional comments (optional):
None.
2.3 Data crawled and scraped from online sources
Were crawlers used by the provider or on behalf of?
- [ ] Yes
- [X] No
WAINUT operated no web crawler for this modification, nor did anyone crawl on WAINUT's behalf. All third-party material entered through published datasets, reported under Section 2.1.
If yes, specify crawler name(s)/identifier(s):
N/A.
Purposes of the crawler(s):
N/A.
General description of crawler behaviour:
N/A for WAINUT's own pipeline. For completeness: the web-derived part of the third-party corpus reaches WAINUT only through that corpus. Its compiler restricted its web extract to pages carrying an explicit CC0 or CC BY statement in the HTML, with sites blocked by robots.txt excluded. WAINUT did not verify that filtering itself.
Period of data collection:
N/A for WAINUT. The third-party corpus is pinned to a revision of 4 May 2026.
Comprehensive description of the type of content and online sources crawled:
N/A.
Type of modality covered:
N/A.
Summary of the most relevant domain names crawled:
N/A. Because no crawler was used by WAINUT or on WAINUT's behalf, the domain-name disclosure does not arise.
Additional comments (optional):
WAINUT's own rule, recorded in its internal records, is that a claim of "no web crawling" may never be made about the model as a whole but only about WAINUT's own layer. The base model was trained on web data by its own publisher, whose Summary covers that.
2.4 User data
Was data from user interactions with the AI model (e.g. user input and prompts) used to train the model?
- [ ] Yes
- [X] No
Was data from user interactions with other services or products of the provider used to train the model?
- [ ] Yes
- [X] No
Type of modality covered:
None.
Additional comments (optional):
WAINUT had no deployed product generating user interactions during the development of this model.
2.5 Synthetic data
Have you used synthetic data to train the model?
- [X] Yes
- [ ] No
Modalities covered:
- [X] Text
- [ ] Image
- [ ] Audio
- [ ] Video
If yes, specify the general-purpose AI model(s) used to generate the synthetic data if available on the market:
swiss-ai/Apertus-70B-Instruct-2509, published under the Apache License 2.0. It generated the answers in the instruction data, on 51 hand-written seed prompts and on passages drawn from the public corpus named in Section 2.1. Generation settings: temperature 0.8, top_p 0.9, max_tokens 1024, seed 1234. In the instruction mix that was actually trained, 47,488 of 51,875 records carry that model as their recorded generator.
That model has the same publisher as the base model and is covered by the same Public Summary of Training Content, at https://huggingface.co/swiss-ai/Apertus-70B-2509/blob/main/Apertus_EU_Public_Summary.pdf. The model card of that model links to the same file. That Summary lists Apertus-70B-Instruct among the models it covers.
Information about other AI models, including provider's own AI model(s) not available on the market, used to generate synthetic data:
The post-training layer of 1,157 records consists of WAINUT's own generations from earlier versions of this model made during its own training, which are not available on the market. No new external source was added at that stage. The model card states: "Daar komt dus geen nieuwe externe bron bij: de vragen komen uit dezelfde verzameling als hierboven en de antwoorden zijn eigen generaties." [So no new external source is added there: the questions come from the same collection as above and the answers are our own generations.]
Those 1,157 records were selected from generation pools of 24,000, 12,000 and 15,400 records. Inside the layer itself the origin splits as 871 records from one generation run, 145 from a second and 141 from a third. No source records which pool a given part of that split came from, so the two sets of figures are not matched up here.
Additional comments (optional):
No prompt left the compute node for an external API during generation. The generator was the model named above and no other. WAINUT's own records of the instruction data carry that as a compliance condition of the generation run.
WAINUT's own standing rule has two parts. The licence ceiling is that synthetic data may come only from models under an Apache 2.0 or MIT licence. Within that ceiling WAINUT limits itself further to two named models, Apertus and OLMo. The same rule forbids outputs of OpenAI, Anthropic or Gemini models outright. It also requires the generator model, the prompts and the filtering to be documented.
The licence of the teacher model does not restrict training on its output. WAINUT's own licence assessment records that training on the output is neither forbidden nor tied to a condition.
2.6 Other sources of data
Have you used other sources of data to train the model?
- [X] Yes
- [ ] No
If yes, provide a narrative description of these data sources and the data:
Two elements fall here.
- Human-written seed prompts. 51 seed prompts were written by hand. They are the starting point for the self-instruction part of the instruction data. They are Dutch text and they were not taken from any external source. What the records of that data carry is the count and the fact that the prompts are hand-written. They name no author, so none is named here.
- Selection in the post-training layer. In its last training step the released model was trained on its own answers that an assessment step had approved. No preference pairs were used.
No person made the selection in that step. In the three families with a checkable answer (classification, extraction and redaction of personal data) a rule-based check decided; that check does not measure whether anything was made up. Outside those three families the assessment step was a panel of pinned judge models of third parties, each hosted at its own provider, with provider pinning set and fallbacks switched off. Two judges assessed each answer and an answer counted as good when both called it good. When they disagreed a third judge decided by a majority of two of three. At most one answer was kept per prompt. An answer was dropped when it ran longer than 1.8 times the median length in characters of the approved answers of the same layer. Where that rule dropped the only approved answer of a prompt, the prompt produced no record at all. The recipe of that step goes to the public repository on GitHub at https://github.com/WAINUTAI/walnoot-8b-instruct, so a reader can check this.
Additional comments (optional):
The instruction data of the released model holds 51,875 records in twenty task categories: summarisation 6,858, instruction following 5,998, reasoning and arithmetic 5,600, government and legal question answering 5,001, text simplification 3,986, Dutch text classification 3,853, refusal on safety grounds 3,438, multiple choice 3,214, rewriting 2,832, English scientific annotation 2,491, classification 2,400, extraction 2,017, English retention 1,500, rewriting into plain Dutch at level B1 660, letter drafting 547, abstention 500, commercial arithmetic 400, identity 250, behavioural layer 247 and redaction of personal data in a supplied text 83. Those twenty counts were measured on 23 September 2026 over the instruction data that was trained.
3. Data processing aspects
3.1 Respect of reservation of rights from text and data mining exception or limitation
Are you a Signatory to the Code of Practice for general-purpose AI models that includes commitments to respect reservations of rights from the TDM exception or limitation?
- [ ] Yes
- [X] No
WAINUT has not signed the Code of Practice. WAINUT follows the Copyright Chapter of that Code as a guideline all the same; its own copyright policy is written to the five measures of that Chapter.
Description of the measures taken to respect reservations of rights before and during data collection, including the opt-out protocols honoured, also those of third parties from whom datasets were obtained:
WAINUT ran no crawler of its own for this modification, so WAINUT took no robots.txt decision of its own. Reservations of rights are honoured one step upstream, in two places.
- In the third-party Dutch corpus. Its web-derived part is a licence-filtered extract that keeps only pages carrying an explicit CC0 or CC BY statement in the HTML, with sites blocked by robots.txt excluded. WAINUT adds no fresh web data of its own.
- In the base model. Its publisher states in its own code of practice: "In addition, data from websites which have recently opted out by specifying at least one of the common AI crawlers, at the time of January 2025, was removed. Crucially, such removals were also applied retroactively in all earlier crawls since 2013, of each corresponding website present in our datasets." That is the publisher's statement about the base model, not a measurement by WAINUT.
Further measures on WAINUT's own layer: every external source in the instruction mix was traced to a primary licence source, the tenth row of the triage being WAINUT's own identity layer, which comes from no external source; no non-commercial licence, no share-alike licence, no GPL and no research-only licence occurs in the mix that was trained; sources whose most specific primary licence forbids commercial use were dropped rather than used.
What WAINUT did not do, kept apart from what the publisher did. WAINUT ran no crawler, so WAINUT honoured no reservation of rights of its own during collection. The material reached WAINUT through the open dataset repository of the publisher, so the reservation of rights that counts for this material is the one honoured by the publisher of the dataset. On the two archive domains behind the subsets discussed in Section 2.1 there is no robots.txt. That is a measured fact about those domains. It is not an argument, because WAINUT took nothing from those sites.
Additional comments (optional):
WAINUT's copyright policy for this model is set out in a separate document, which follows the five measures of the Copyright Chapter of the Code of Practice. It is published at https://walnoot.ai/en/transparantie beside this Summary. Rightsholders, downstream providers and anyone with a complaint or a question can reach WAINUT at support@walnoot.ai.
3.2 Removal of illegal content
General description of measures taken:
- Rejection of candidate passages with personal data. A gate ran over the verbatim passages from the corpus that were set aside as candidates for document tasks in the instruction data. It rejected 16,776 of those passages because personal data was found in them. They were left out rather than reworked. A re-run of the gate over the delivered file returned zero hits on 1,600 passages. The gate covered only those candidate passages; the rest of the training data did not go through it.
- Malware-flagged files excluded. Eight parquet files in one source subset that a virus scanner had flagged were excluded by an explicit exclusion list.
- Personal data detection. An eighteen-detector filter exists, combining patterns with arithmetic validation, for example the eleven-test for Dutch social security numbers, mod 97 for IBANs and Luhn for card numbers. Recall per class has not been measured and cannot be measured with this set-up, because six of the ten classes have an empty denominator. What was measured is the recall on the classes that do occur in a stratified sample of 300 records, using the same pinned detector set: zero for private-person names, zero for indirectly identifying data, plus zero for the removal rule against all nine real cases in that sample. The planned second layer never ran; WAINUT's internal records name it as a Dutch named-entity model and advise against installing it. It did not run over the training data as a whole. The measurement that does exist is an after-the-fact count on the post-training layer of 1,157 records: eleven of eighteen detectors fire, together 702 hits, split as 528 hits on the input side and 174 on the output side. The seven detectors that return zero were shown to fire on synthetic test cases, so the zero is not a silent detector. The detectors found no Dutch social security number, no credit card number, no date of birth in context and no bank account number that passes the arithmetic check. Addresses, postcodes, vehicle registrations, telephone numbers and e-mail addresses were found. What was measured on this data itself is this: no e-mail address in it combines a person-shaped local part with an existing domain. The judgment that the other kinds are predominantly synthetic or point to a reserved domain rests on a hand reading of comparable data from an earlier round and not on a hand reading of this data.
- Two data families work with personal data by design because they teach the model to mark or remove such data. That their values come from reserved ranges is an indication and not a test, which is how the model card records it: the arithmetic check on bank account numbers fires zero times in both families while account-number-shaped values do occur in them. A hand reading of these two families in particular has not been done. There is no exemption from masking for these two families. The question does not arise, because the training data as a whole did not go through a personal data filter, as Section 1.3 records.
- Public case law is not reworked, because the source itself already applies its own protection. Person names in a professional context were kept. The ground recorded for that choice is the legal basis and the balancing test, not the fact that the source is public.
- What the publisher of the corpus states itself about one of the two archive subsets. In its collection metadata, retrieved on 22 September 2026, the publisher writes about the Noord-Hollands Archief subset: "Names of citizens (e.g., those filing complaints) and addresses appear in the documents. These are part of the original municipal records." That is a statement by the publisher about its own collection and not a measurement by WAINUT. It is reported here because it bears on the point above: the training data as a whole did not pass a personal data filter, so this material entered the corpus as the publisher describes it.
- No dedicated detection for the heaviest categories of illegal material. WAINUT ran no dedicated detection for child sexual abuse material or for non-consensual intimate imagery over its own layer. The training data is text without images.
3.3 Other information (optional)
Decontamination against the public evaluation battery. The reference set of the two measurements of July 2026 is 35 EuroEval dataset splits holding 29,273 records, indexed as 2,980,615 unique n-gram fingerprints. Those 35 splits are not all test splits: WAINUT's own measurement records count train and validation splits among them, so they come from roughly thirteen distinct datasets. The model card names the version behind that reference set: "35 bestanden uit EuroEval 17.6.0, samen 29.273 records, elk met een eigen controlegetal." [35 files from EuroEval 17.6.0, together 29,273 records, each with its own checksum.] That version comes from the model card. WAINUT's own measurement records name only the version of the measuring tool.
There are three measurements and they are not about the same material. The first pass, of 22 July 2026, covered an earlier and smaller build of the corpus of the continued pretraining and the instruction data of that moment, 238,971 documents together. At thirteen words in a row it touched 16 of the 29,273 records, 0.0547 percent, which against a fail threshold of zero is a rejected outcome. The depth pass at sixteen, twenty and twenty-five words left one real hit of those sixteen: a widely circulated fixed prayer text in one test item. The other fifteen were fixed-phrase false positives that did not survive the depth pass. That first pass has not been run again on the build of the corpus on which the model was trained. The second measurement, of 30 July 2026, covered the two source files of WAINUT's own source of synthetic chat pairs for preserving the formatting of text in the corpus (see Section 2.1). It comes out at zero at thirteen words, measured per source, for both sources. The binding record of that second measurement states that no overlap was found: none of the 29,273 evaluation records shares a thirteen-word sequence with those two files. An eight-word check returns 723 records for one source and 591 for the other. That check calls itself a gate in the same words as the thirteen-word one. It runs with a fail threshold of 100 percent of contaminated records per dataset, so in practice it reports rather than blocks; the blocking gate is the thirteen-word check, whose threshold is 0 percent.
The third measurement, of 12 September 2026, covered the files on which the released model was actually trained: the instruction data of 51,875 records and the post-training data of 1,157 records. It ran after the training, against the ten Dutch tasks of EuroEval on which the model was measured: 28 splits with 42,925 items together, each measured once on the question and once on the gold answer. Its gate lies at twenty words. In the instruction data two items from the test split of squad-nl were touched at thirteen words, on the question side. The shared stretch covers at most 1.27 percent of such an item and is gone at sixteen words, so it is a fixed phrase. Nothing was touched at twenty words or more. In the post-training data nothing was touched at thirteen words or more. The outcome is clean. An eight-word audit touched 659 items in the instruction data and 33 in the post-training data. A summary of that measurement is published in the folder decontaminatie of the public repository on GitHub at https://github.com/WAINUTAI/walnoot-8b-instruct.
Model quality safeguards that are not data processing. The model card records four numbered limitations, among them that the model must not be used to remove personal data from documents. That the model has not been tested by a structured red-team round the model card records elsewhere, with its results rather than in that numbered list.
Legal basis under data protection law. WAINUT relies on legitimate interest under Article 6(1)(f) GDPR as the basis for the processing. The balancing for it is recorded internally along the cumulative three-step test of EDPB Opinion 28/2024. That record dates from July 2026. It has not been updated for this release and no legal adviser has reviewed it. The model card adds: "Hoeveel de detectie mist, is alleen in een steekproef gemeten. In 300 records uit onze eigen databestanden, waarvan 90 uit de instructiedata van dit model, vond dezelfde detectorset geen van de negen echte gevallen van persoonsgegevens. Op de nabewerkingsdata zelf en per soort is het niet gemeten. Wie dit model in een omgeving met persoonsgegevens inzet, blijft zelf verwerkingsverantwoordelijke en kan zich niet op deze meting beroepen." [How much the detection misses has been measured only in a sample. In 300 records from our own data files, 90 of them from the instruction data of this model, the same detector set found none of the nine real cases of personal data. On the post-training data itself and per class it has not been measured. Anyone who deploys this model in an environment with personal data remains the data controller and cannot rely on this measurement.]
Compute acknowledgement, required verbatim in every publication about this model:
Required acknowledgement, verbatim:
"We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-984 access to MareNostrum5 ACC hosted by BSC, Spain"