The tag says apache-2.0
A repository card declares apache-2.0. The model was fine-tuned from a Llama base, on a dataset generated by calling a commercial API, using a merge that pulled in a fourth model. Three separate sets of obligations survived that pipeline, and none of them is visible in the tag.
A licence field describes what the publisher of one repository asserts about one repository. It does not compute the chain, and no metadata field anywhere on the Hub does.
Four licences, not one
The obligations attached to a derived model are the union of at least four things: the licence on the base model weights, the licence on every training dataset, the terms of service of any model whose outputs became training data, and the licence of every constituent in a merge or adapter stack.
The result is governed by the intersection of what those permit — which is to say, by the most restrictive of them. One non-commercial dataset in the pipeline makes the output non-commercial regardless of what the base licence allowed.
Llama: a threshold, a prefix, and an attribution
The Llama community licence carries three obligations that regularly surprise people who read only the word community.
First, a user threshold. If, on the version release date, the monthly active users of the products or services made available by or for the licensee or its affiliates exceeded 700 million in the preceding calendar month, the licensee must request a separate licence from Meta. This is not a usage cap; it is a rule about who you are.
Second, naming. If Llama Materials, or any outputs or results of them, are used to create, train, fine-tune or otherwise improve an AI model that is distributed or made available, the licensee must include Llama at the beginning of that model's name. Note that the clause covers outputs, so it reaches synthetic data generated by a Llama model, not only weight derivatives.
Third, attribution. The licensee must prominently display Built with Llama on a related website, user interface, blog post, about page or product documentation.
None of these is onerous. All of them are frequently unmet by repositories that describe themselves as openly licensed.
Gemma: the restrictions run with the weights
The Gemma terms take a different approach: they make the restrictions travel downstream by contract. Distribution of a model derivative requires including the use restrictions as an enforceable provision in any agreement with the recipient, and providing that recipient a copy of the agreement.
Distributions outside a hosted service must also be accompanied by a notice file containing specific text stating that Gemma is provided under and subject to the Gemma Terms of Use.
The practical effect is that a downstream user of a Gemma derivative is bound by the Gemma use restrictions whatever the derivative's own README says. A permissive tag on a Gemma derivative does not release anybody from anything.
Synthetic data carries the terms of the model that produced it
Stanford Alpaca is the canonical example because its authors documented the reasoning. The dataset is released under CC BY-NC 4.0 — non-commercial — and the authors state that models trained on it should not be used outside research.
There are two independent reasons for that restriction, and both are inherited rather than chosen. The base model, LLaMA, was released under a non-commercial licence. And the instruction data was generated with OpenAI's text-davinci-003, whose terms of use prohibited developing models that compete with OpenAI.
This is the general case. A dataset generated by calling any commercial API inherits that API's terms whether or not the dataset file mentions them, and whether or not the person who published the dataset was aware of them.
An undeclared licence is not a permissive one
A substantial number of repositories declare no licence at all. Absence of a grant is not a grant. Without a licence, default copyright applies to the weights, and the safe reading is that you have no permission to redistribute or to build on them commercially.
The catalogue on this site records those repositories as undeclared, with the reason stated, rather than defaulting them to permissive or leaving the field blank. A blank field reads as an oversight; an explicit undetermined reads as a finding.
EU AI Act Article 53, and what open release does not exempt
Providers of general-purpose AI models carry four duties under Article 53(1): (a) draw up and keep current technical documentation of the model including its training and testing process, at minimum the information in Annex XI; (b) draw up and make available information and documentation to providers of AI systems intending to integrate the model; (c) put in place a policy to comply with Union copyright law, including identifying and complying with rights reservations expressed under Article 4(3) of Directive (EU) 2019/790; and (d) draw up and make publicly available a sufficiently detailed summary about the content used for training, according to a template provided by the AI Office.
Article 53(2) exempts models released under a free and open-source licence that allows access, usage, modification and distribution, and whose parameters, architecture information and usage information are public — but the exemption applies to points (a) and (b) only. It does not apply at all to models with systemic risk.
So open release removes the two documentation duties and leaves the copyright policy and the public training-content summary in place. A team that assumes an open licence discharges Article 53 has misread the paragraph that follows the one they read.
The audit that actually resolves the chain
- Name the base model at a commit SHA, and read the licence file in that repository rather than the card metadata.
- List every training dataset and its licence, including datasets embedded in a mixture you did not assemble.
- For every dataset generated by a model, record which model and check that model's terms of service, not the dataset's card.
- List every merge constituent and adapter, and repeat the first three steps for each.
- Take the intersection. Where any input is non-commercial or research-only, the output is too.
- Check the naming, attribution and downstream-enforceability duties of every licence in the chain against what you are actually about to publish.
- Record the Article 53 position: whether the model is general-purpose, whether the open-source exemption applies, and how the copyright policy and training-content summary will be published.
Not legal advice
This lesson reads published licence text and one regulation. It is not legal advice, it does not account for jurisdiction, and it cannot see the contracts you have already signed. Use it to know which questions to take to counsel.