Skip to main content

GPAI documents

Copyright policy

How we deal with copyright and with reservations of rights under text and data mining.

Model
Walnoot 8B Instruct 1.0
Issued
28 September 2026

Back to GPAI transparency

Document dated 28 September 2026. Model version 1.0, released on 28 September 2026.

Published voluntarily by WAINUT; see the statement in this document.

This document has not yet been reviewed by a legal adviser.

Quotations from EU documents are reproduced verbatim in English.

A policy to comply with Union law on copyright and related rights, written along the five Measures of the Copyright Chapter of the Code of Practice for general-purpose AI models.


The provider, the model and this document

Provider: WAiNuT B.V. (registered legal name), trading as WAINUT, a private limited company registered with the Netherlands Chamber of Commerce (KVK) under number 93560915, Rivium Westlaan 48, 2909LD Capelle aan den IJssel, the Netherlands. Point of contact: support@walnoot.ai. Website: https://walnoot.ai.

fieldvalue
Model this policy coversWalnoot 8B Instruct, version 1.0
Model repositoryhttps://huggingface.co/WAINUT/walnoot-8b-instruct
Public repository on GitHubhttps://github.com/WAINUTAI/walnoot-8b-instruct
Issued and placed on the Union market28 September 2026
This documentversion of 28 September 2026
Published athttps://walnoot.ai/en/transparantie in English, https://walnoot.ai/transparantie in Dutch

The model repository and the public repository on GitHub go live with the release on 28 September 2026.

The legal name is written above in the spelling that the commercial register carries. Everywhere else in this document the trade name WAINUT is used.

Status, scope and conventions

Status. By its own measurement WAINUT is not a provider of a general-purpose AI model within the meaning of Regulation (EU) 2024/1689 (the AI Act), because the training compute of its modification stays far below the indicative threshold of the Commission Guidelines. That measurement is recorded in WAINUT's internal records and is not published. WAINUT is not a Signatory to the Code of Practice and follows the Code as a guideline. This policy is published voluntarily on the transparency page of WAINUT, in English at https://walnoot.ai/en/transparantie and in Dutch at https://walnoot.ai/transparantie. The model card carries a reference to that page. The public summary of the training content and the column of the Model Documentation Form for downstream providers are published on that same page. The Model Documentation Form as a whole is not published: it is kept and handed to the AI Office or to the national authority on request.

Two things make that choice deliberate rather than decorative. First, the Commission states that a free and open-source release does not remove the copyright obligation: "providers of general-purpose AI models that meet the conditions referred to in Section 4.2 are not exempted from the obligation laid down in Article 53(1), point (c), AI Act to put in place a policy to comply with Union copyright law. This obligation includes identifying and complying with a reservation of rights expressed under Article 4(3) of Directive (EU) 2019/790 of the European Parliament and of the Council." Second, openness about provenance is the point of this project, so a reader should be able to see what WAINUT did, in a single document.

Scope. This policy covers WAINUT's own modification of the base model. Where the obligations of a provider would apply, they are limited to the modification: "the copyright policy required by Article 53(1), point (c), AI Act and the summary of the content used for training required by Article 53(1), point (d), AI Act are limited to the data used as part of the modification." For the base model, its own publisher's policy and summary apply.

This is a single document. Measure 1.1 asks for the policy to be described in one document incorporating the Measures of the Chapter. This is that document.

Conventions. Nothing WAINUT did not do is claimed here. No internal paths, host names, keys or storage addresses appear in this document. No prompt, record or answer text from the training data appears in this document. Dutch quotations are reproduced verbatim; the English rendering in square brackets after each of them is WAINUT's own translation and is not authoritative.


Measure 1.1 Draw up, keep up-to-date and implement a copyright policy

The Measure reads: "Signatories will draw up, keep up-to-date and implement a policy to comply with Union law on copyright and related rights for all general-purpose AI models they place on the Union market. Signatories commit to describe that policy in a single document incorporating the Measures set out in this Chapter. Signatories will assign responsibilities within their organisation for the implementation and overseeing of this policy."

What WAINUT did. WAINUT wrote down a data standard: lawfully obtained material that respects reservations of rights, with the publicly released CC BY corpus of the GPT-NL consortium named among the permitted sources. The standard is recorded in an internal legal checklist. That checklist is a working document and it records no compliance of its own, so it is not offered here as evidence that the standard was worked to. What was published with the model is what shows how the standard was applied: a model card that states the provenance per source, a NOTICE file carrying the attribution list, a provenance manifest per subset and a source-level licence triage over the instruction mix that was actually trained.

Three rules from that standard shaped the data:

  1. Every licence must be traced to a primary source, meaning the licence file or dataset card of the source set itself, or the appendix of that set's own paper. Not a paper table about another set and not an intermediate aggregator.
  2. WAINUT's own standard excludes share-alike and non-commercial licensed material, which is why Wikipedia is not among its permitted sources. That rule stands, also now that another Dutch research group counts CC BY-SA as permissibly licensed. It costs this release nothing: a count over the licence field of the instruction mix that was trained returns ten licence rows and no share-alike row among them.
  3. Synthetic data may be generated only with models under an Apache 2.0 or MIT licence. Within that ceiling WAINUT limits itself further to two named models, Apertus and OLMo. Outputs of OpenAI, Anthropic or Gemini models are not used. The generator model, the prompts and the filtering have to be documented.

Keeping it up to date. This policy is reviewed once a year and at every new model release by WAINUT. A revision is approved by the owner of WAINUT.

Responsibilities. The management of WAINUT carries the responsibility for implementing and overseeing this policy. Everything that touches this policy goes to support@walnoot.ai.


Measure 1.2 Reproduce and extract only lawfully accessible copyright-protected content when crawling the World Wide Web

The Measure commits signatories "not to circumvent effective technological measures as defined in Article 6(3) of Directive 2001/29/EC that are designed to prevent or restrict unauthorised acts in respect of works and other protected subject matter, in particular by respecting any technological denial or restriction of access imposed by subscription models or paywalls". It further commits them "to exclude from their web-crawling websites that make available to the public content and which are, at the time of web-crawling, recognised as persistently and repeatedly infringing copyright and related rights on a commercial scale by courts or public authorities in the European Union and the European Economic Area."

What WAINUT did.

  1. WAINUT operated no web crawler of its own. There is therefore no crawl of WAINUT's own in which a technological measure could have been circumvented and no crawl of WAINUT's own from which a site would have had to be excluded. WAINUT declares as well that nobody crawled on WAINUT's behalf: no crawl was commissioned and no web page was fetched to order. The material reached WAINUT through the open dataset repository of the publisher of the corpus and through published datasets. That declaration is the source for this point; there is no other.
  2. All third-party material was obtained from published datasets through their own distribution channels. The Dutch layer comes from one public, ungated repository, the GPT-NL Public Corpus, released by the consortium that built it under CC BY 4.0, downloaded at a pinned revision. The scientific branch comes from fourteen public annotation datasets, each of whose licences was traced to a primary source. Those fourteen did not arrive the same way: their primary sources sit on the distribution channels of the individual sets rather than on one pinned repository revision. One of the fourteen, AnatEM, carries CC BY-SA 3.0 on its primary source. Its 168 records do not reach the instruction mix that was trained; see the note under Attribution below.
  3. No paywalled, subscription-restricted or otherwise access-restricted source was used. No access control was bypassed, because no source used required one.
  4. Web-derived material is present only one step upstream. Inside the third-party corpus the web extract keeps only pages carrying an explicit CC0 or CC BY statement in the HTML, with sites blocked by robots.txt excluded. WAINUT relies on that filtering by the corpus compiler; WAINUT did not re-verify it and does not claim to have done so.
  5. A source-level licence gate was applied to the instruction data. It cost records. Of 9,722 candidate records 4,421 passed the primary-source rule; 2,659 remained after all the later quality gates. The tightening of the rule itself dropped three source sets with 281 records between them. One of those three is the case that prompted the rule: no primary source permits commercial use, while the most specific source available, the intermediate resource from which the corpus itself draws that set, states a non-commercial licence where the table in the corpus paper claimed CC BY. At the licence gate as a whole more fell away than those three: five named source sets, plus the sixteen source sets for which that table fixes no licence at all so that no primary source was sought. Those sixteen hold 4,749 records between them.
  6. Government texts. Dutch statutes, decisions and court judgments are free of copyright under Articles 11 and 15b of the Dutch Copyright Act, which is the basis on which the public-sector and case-law material in the corpus is used.
  7. An own tightening above the source. Eight parquet files in one subset that a virus scanner had flagged were excluded by an explicit exclusion list.

What WAINUT does not claim. WAINUT never claims "no web crawling" for the model as a whole. The claim holds only for WAINUT's own layer. The base model was trained on web data by its own publisher, whose own policy and summary cover that.


Measure 1.3 Identify and comply with rights reservations when crawling the World Wide Web

The Measure commits signatories "to employ web-crawlers that read and follow instructions expressed in accordance with the Robot Exclusion Protocol (robots.txt), as specified in the Internet Engineering Task Force (IETF) Request for Comments No. 9309, and any subsequent version of this Protocol". Its fourth paragraph asks signatories to make public information about the crawlers employed and their robots.txt features, with a means for affected rightsholders to be notified automatically when that information is updated.

What WAINUT did.

  1. WAINUT employs no web crawler, so WAINUT takes no robots.txt decision of its own. There is no crawler name, no user agent and no crawl schedule to publish, because none exists. The obligation to publish crawler information has no subject on WAINUT's side.
  2. Reservations of rights are honoured upstream, in two places, by parties that did crawl. In the third-party Dutch corpus, through the licence-filtered web extract described under Measure 1.2, which excludes sites blocked by robots.txt. In the base model, whose publisher states in its own code of practice: "In addition, data from websites which have recently opted out by specifying at least one of the common AI crawlers, at the time of January 2025, was removed. Crucially, such removals were also applied retroactively in all earlier crawls since 2013, of each corresponding website present in our datasets." That is the publisher's own statement about its own model. WAINUT reproduces it here as the reason the upstream layer is acceptable to WAINUT; WAINUT did not verify it.
  3. WAINUT adds no fresh web data. Nothing was scraped after the corpus revision was pinned.
  4. What WAINUT fetched and what others fetched are two different things. Neither of the two archive domains behind the two subsets described below carries a robots.txt and no reservation against text and data mining was found on the pages that were read there. That is a measured fact and it is not an argument for anything, because WAINUT fetched nothing from those sites: the material reached WAINUT through the open dataset repository of the publisher of the corpus. The reservation of rights that counts for this material is therefore the one that the publisher of that dataset had to identify and comply with.

What is therefore not in place. WAINUT publishes no crawler information and offers no web feed notifying rightsholders of changes to crawler behaviour, because WAINUT operates no crawler. If WAINUT ever crawls, this Measure has to be implemented in full before that crawl starts.


Measure 1.4 Mitigate the risk of copyright-infringing outputs

The Measure commits signatories "to implement appropriate and proportionate technical safeguards to prevent their models from generating outputs that reproduce training content protected by Union law on copyright and related rights in an infringing manner". It further commits them "to prohibit copyright-infringing uses of a model in their acceptable use policy, terms and conditions, or other equivalent documents, or in case of general-purpose AI models released under free and open source licenses to alert users to the prohibition of copyright infringing uses of the model in the documentation accompanying the model without prejudice to the free and open source nature of the license."

What WAINUT did.

  1. No training data is redistributed. The model card states: "Er wordt geen trainingsdata gepubliceerd." [No training data is published.] WAINUT publishes weights, documentation and recipe, not corpora. That removes the most direct route by which protected material could be re-published.
  2. Deduplication during the corpus build. Near-duplicate removal is the largest single drop reason in the corpus build at 377,787 documents, out of 806,473 drops over nineteen reasons. This was done for data quality. It reduces repeated exposure to identical passages, which is a known driver of verbatim memorisation.
  3. The instruction layer is largely generated rather than copied. In the mix that was trained, 47,488 of 51,875 records carry an openly licensed teacher model as the recorded source of the answer text. The post-training layer of 1,157 records adds no new external source: the answers are WAINUT's own generations, selected by an assessment step. The questions come from the same collection as the instruction layer, which is accounted for above. The model card adds that for part of that layer the origin cannot be positively established and that this part is counted as non-synthetic, on the strict side.
  4. The base model carries its publisher's own mitigation. Its publisher describes a training technique that, in its own words, "avoids verbatim memorization of text sequences longer than 50 tokens". WAINUT did not re-measure that property after continued pretraining.

What WAINUT did not do, stated plainly. WAINUT implemented no technical safeguard of its own against copyright-infringing outputs and ran no memorisation test on the released weights. The n-gram decontamination that WAINUT did run, at thirteen words against the public evaluation battery, was built to keep evaluation honest, not to measure reproduction of protected works. It must not be presented as such.

The output filter of the base model publisher. The acceptable use policy of that publisher advises applying a file of hash values as an output filter. The publisher writes on the card of the base model that no such filter is provided at the moment, in its own words: "Currently no output filter is provided." The model card of Walnoot 8B Instruct says the same. WAINUT is putting this question to the publisher in September 2026 and will record the answer here. No answer had been received when this document went out.

What WAINUT will do. No memorisation measurement is made on the released weights before publication. WAINUT carries that measurement out after the release and records the result here. The model card carries a sentence alerting users that copyright-infringing uses of the model are prohibited, which is what this Measure asks of a free and open-source release.


Measure 1.5 Designate a point of contact and enable the lodging of complaints

The Measure reads: "Signatories commit to designate a point of contact for electronic communication with affected rightsholders and provide easily accessible information about it." Its second paragraph asks for an electronic complaints mechanism, with complaints acted on "in a diligent, non-arbitrary manner and within a reasonable time, unless a complaint is manifestly unfounded or the Signatory has already responded to an identical complaint by the same rightsholder."

The point of contact. WAINUT designates support@walnoot.ai as the point of contact for electronic communication with affected rightsholders and for complaints about material in the training data. That decision was taken on 22 September 2026. The address stands in the model card of Walnoot 8B Instruct and is published on https://walnoot.ai/en/transparantie as well. The model card names that same address for three cases, one of which is this one: a rightsholder who wants to report something about material in the training data, a report about personal data and an error in the card itself.

The general business details, unchanged and public: WAINUT, Rivium Westlaan 48, 2909LD Capelle aan den IJssel, the Netherlands, commercial register number 93560915, website https://walnoot.ai. For anything that touches this policy, support@walnoot.ai is the address. No second e-mail address and no second domain are named here, so that a rightsholder has one address to write to and one place to look.

The complaints mechanism. A rightsholder who wants to report something about material in the training data writes to support@walnoot.ai. There is no separate electronic form and no second address, so that one address serves for everything this policy covers. WAINUT confirms receipt of a complaint within five working days and answers on the substance within thirty days. A complaint is handled inside WAINUT through that same address. The postal address is the business address above.

What WAINUT promises a rightsholder. On a complaint that proves well founded, WAINUT removes the source concerned from the training mix of a next version of the model. WAINUT promises no retraining of a model that is already released and no withdrawal of one, because no list exists of the documents that were actually used, so such a promise could not be kept. WAINUT runs no output-side copyright filter in real time either; that is stated under Measure 1.4 above.


Attribution

Some of the sources used require attribution. The NOTICE file shipped with the model carries the list, with rightsholder, place of publication and licence per source. It names the GPT-NL Public Corpus, Chia and Qasper under CC BY 4.0, BioASQ under CC BY 2.5 and JNLPBA under CC BY 3.0 with clause 4 of the GENIA Project License. It names SciRIFF, the collection from which the scientific branch comes, under ODC-By. It names SciTLDR and COVID-QA under Apache 2.0, QASA, PubMedQA and MatSci Text Corpus under MIT and LINNAEUS under BSD; those licences ask for attribution as well. Sources whose licence asks no attribution are listed too, for completeness: MedMentions and NLM-Chem under CC0 1.0 and NCBI Disease in the public domain. The teacher model used for the synthetic answers is named there as well, with its Apache License 2.0.

The licence composition of the instruction mix that was actually trained holds ten rows in the triage. Eight of those ten come from external sources and all eight permit commercial use. The ninth, Apache-2.0-clean, is not a licence of an external source but a label WAINUT derives for the generated material of the mix. The records that rest on a literal passage from the corpus carry that label as well; the passage itself falls under CC BY 4.0, the licence of the corpus. The corpus is attributed in the NOTICE file. The tenth is WAINUT's own identity layer, which comes from no external source at all. That is the count the model card carries. No non-commercial licence, no share-alike licence, no GPL and no research-only licence occurs in that triage.

One source set belongs beside that last sentence. One set in the scientific branch, AnatEM, is licensed under CC BY-SA 3.0 and delivered 168 records. Those 168 records do not reach the mix that was trained. The branch enters that mix with 2,491 records, which is the 2,659 delivered records without the CC BY-SA 3.0 row. The eight other licence rows match record for record: CC BY 2.5 812, Apache 2.0 516, MIT 412, CC0 1.0 281, CC BY 4.0 279, public domain 171, CC BY 3.0 16 and BSD 4, together 2,491. A count over the licence field of that mix returns no share-alike row at all. No attribution or share-alike obligation from that source set therefore rests on this release. That was settled on 23 September 2026 by that count. The NOTICE list above names thirteen annotation sets and correctly does not name AnatEM.


The two corpus subsets with the licence value "unknown"

WAINUT's own rule is that a licence value outside a fixed set stops the pipeline for investigation, with "unknown" named in it. That rule was run for these two subsets on 22 September 2026. This is the outcome, recorded here on 23 September 2026.

What carries the value. At the revision of the corpus on which the continued pretraining ran, two of the 28 corpus subsets that were used carry the licence value "unknown": a subset of 7,995 documents from Noord-Hollands Archief and a subset of 73,174 documents from Het Utrechts Archief. That value stands in WAINUT's own provenance manifest, the file herkomst/provenance-manifest.json that goes to the public repository on GitHub at https://github.com/WAINUTAI/walnoot-8b-instruct, in the folder herkomst. It was written there by the script that built the corpus, from the licence column per document in the parquet files of the corpus as they stand at that revision. The per-subset statistics files that the publisher of the corpus ships in its dataset repository carry the same value at that revision. That revision is 1e17d21afa6518939523fbdcc1fbb3167b671725. The other 26 subsets each carry an explicit value. On 7 September 2026 at 11:34 UTC, after that revision, the publisher set both subsets to "public-domain" in a new commit, revision 1b4408c2e78fced22626123c1e8ad840d5183f12. That commit changed the two statistics files and, according to its description, the licence column of every document in both subsets.

What the two parts weigh. Together they are 81,169 documents, 117,179,121 characters and 37,011,724 estimated tokens. Over all 28 subsets that is 0.43 percent of the estimated tokens and 0.41 percent of the characters.

What the publisher says. In its own collection metadata the publisher records both archives as public-domain. That collection metadata is the ground for reading these two parts as public domain; no other ground is claimed for them here. The file is 452,087 bytes with sha256 46fa6d006f80cd482d8a950bb1c26c55d86d48463832457a29b4cb7842ffc88e, retrieved on 22 September

  1. That file can change. On its dataset card the same publisher puts both archives in the

category Archive/Public Domain and sets the corpus as a whole under CC BY 4.0. The values that card documents for the licence field per document are public domain, CC-0 and CC-by; "unknown" is not among them.

What the two archives say. Both publish their open data under CC0.

Noord-Hollands Archief, on its page with the open data, fetched fresh on 22 September 2026 with no retrieval date recorded per page: "Personen en organisaties kunnen de open data van het Noord-Hollands Archief vrijelijk gebruiken, al dan niet voor commerciële doeleinden. Er geldt de licentie CC0 (no rights reserved), dat wil zeggen iedereen mag de gegevens vrij gebruiken zonder toestemming te vragen. Bronvermelding is niet nodig, maar wordt wel op prijs gesteld." [Persons and organisations may freely use the open data of the Noord-Hollands Archief, whether or not for commercial purposes. The CC0 licence applies (no rights reserved), that is to say anyone may use the data freely without asking permission. Attribution is not required but is appreciated.]

Het Utrechts Archief, on its page with the open data, fetched fresh in the same way and likewise with no retrieval date recorded per page: "Personen en organisaties kunnen de open data vrijelijk gebruiken, al dan niet voor commerciële doeleinden. Wel zijn er enkele spelregels. Voor de nu vrijgegeven gegevens geldt de licentie CC0." [Persons and organisations may freely use the open data, whether or not for commercial purposes. There are, however, a few ground rules. For the data released so far the CC0 licence applies.]

Both statements are about the open data of those archives, that is about their finding aids, not in so many words about the text of the documents themselves. That is the limit of what they prove.

What WAINUT could not establish. Where the value "unknown" came from. In the description of its commit of 7 September 2026 the publisher writes that the value looks like a gap in the metadata rather than an intentional value. Neither the dataset card nor the summary of the article about the corpus says whether the publisher took the value over from the archive or filled it in for want of a field. Three facts weigh against reading that value as a restriction. The publisher records both subsets as public-domain in its own collection metadata. Both archives place their open data under CC0. And what was used is text from character recognition and not image material, so no separate right in a scan or a photograph comes into play: the Utrecht part is described by the publisher as collections more than a hundred years old, with a temporal coverage given as entirely before 1950. The Noord-Holland part consists of municipal council minutes of one municipality, digitised around 2015 through character recognition.

The outcome. Both parts stay in the training data. This policy works at source level, the corpus stands at that level as one source under CC BY 4.0 and the publisher counts both parts as public-domain in its own collection metadata. Two things belong beside that outcome and WAINUT states them itself rather than leaving them to be found. WAINUT did not walk these two subsets at the source itself. And the model has already been trained while no list exists of the documents that were actually used, so these two parts cannot be taken back out of the released weights. WAINUT reads what remains as a gap in the metadata per document at the revision on which the continued pretraining ran. The publisher has since set that value to public-domain and writes that the earlier value looks like a gap in the metadata. WAINUT does not see that gap as a bar. That reading rests on what the publisher and the two archives state as set out above. No legal adviser has reviewed it.


Open items

Three things are open. They are written here in full rather than left to be found.

  1. No legal adviser has read this document. Until one has, it stays a draft.
  2. The publisher of the base model has not answered the question about the output filter. WAINUT is putting this question to the publisher in September 2026 and will record the answer here.
  3. No memorisation measurement has been made on the released weights. WAINUT carries that measurement out after the release and records the result here.

Compute acknowledgement, required verbatim in every publication about this model

Required acknowledgement, verbatim:

"We acknowledge EuroHPC JU for awarding the project ID EHPC-AIF-2026PG01-984 access to MareNostrum5 ACC hosted by BSC, Spain"