CATALOGUE RECORD / HUB-DERIVED
trl-internal-testing
tiny-Qwen3MoeForCausalLM
Revision 6db57163fa56fa0e00d4b820169908ac0d3a219d
TENSOR ACCOUNTING
Where the bytes are.
Summed from the safetensors index, one row per dtype. A parameter count alone cannot produce this figure, because a checkpoint may mix widths.
| Dtype | Parameters | Bytes each | Bytes | Share |
|---|---|---|---|---|
BF16 | 2,603,624 | 2 | 5 MB | 100.0% |
EVERY FIELD, WITH ITS ORIGIN
Sourced or undetermined. Never assumed.
Each value below names the exact API field it was computed from. Where the Hub does not establish a value, the reason is shown instead of a plausible default.
- Architecture
- qwen3_moe
config.model_type - Model classes
- Qwen3MoeForCausalLM
config.architectures - Routed experts
- 4
config.num_experts - Experts per token
- 2
config.num_experts_per_tok - Shared experts
- UndeterminedThe published config declares no always-on shared experts.
- Routing sparsity
- 2.0× (1 of every 2.0 experts)
config.num_experts / config.num_experts_per_tok - Total parameters
- 2,603,624
safetensors.total - Checkpoint bytes
- 5,207,248 (5 MB)
safetensors.parameters - Ships below 16-bit
- No
safetensors.parameters - Quantisation method
- UndeterminedThe repository declares no quantization method, which normally means unquantized weights.
- Trained context
- UndeterminedThe config summary omits max_position_embeddings. Trained context is a model-card claim, not a derivable fact.
- Declared licence
- UndeterminedThe repository declares no license in its card metadata or tags. Absence of a declared license is not permission to use the weights.
- Base model
- UndeterminedThe repository declares no base model.
- Library
- transformers
library_name - Files in repository
- 12
siblings - Last modified
- 2026-05-07
lastModified
LICENCE POSTURE
Undeclared
The repository declares no licence. Absence of a licence is not permission: without a grant, default copyright applies to the weights.
This is a reading of a metadata field, not legal advice, and it does not account for the licences of upstream models or training data.
WEIGHT RESIDENCY FLOOR
The count below which it cannot fit.
Ceiling of checkpoint bytes over advertised accelerator memory. This is a lower bound on accelerator count for weights alone — KV cache, activations and runtime overhead all sit on top, so a real deployment needs more.