CATALOGUE RECORD / HUB-DERIVED

trl-internal-testing

tiny-Qwen3MoeForCausalLM

Revision 6db57163fa56fa0e00d4b820169908ac0d3a219d

3MTOTAL PARAMETERS
5 MBCHECKPOINT BYTES
4 / 2EXPERTS / PER TOKEN
716KDOWNLOADS

TENSOR ACCOUNTING

Where the bytes are.

Summed from the safetensors index, one row per dtype. A parameter count alone cannot produce this figure, because a checkpoint may mix widths.

DtypeParametersBytes eachBytesShare
BF162,603,62425 MB100.0%

EVERY FIELD, WITH ITS ORIGIN

Sourced or undetermined. Never assumed.

Each value below names the exact API field it was computed from. Where the Hub does not establish a value, the reason is shown instead of a plausible default.

Architecture
qwen3_moeconfig.model_type
Model classes
Qwen3MoeForCausalLMconfig.architectures
Routed experts
4config.num_experts
Experts per token
2config.num_experts_per_tok
Shared experts
UndeterminedThe published config declares no always-on shared experts.
Routing sparsity
2.0× (1 of every 2.0 experts)config.num_experts / config.num_experts_per_tok
Total parameters
2,603,624safetensors.total
Checkpoint bytes
5,207,248 (5 MB)safetensors.parameters
Ships below 16-bit
Nosafetensors.parameters
Quantisation method
UndeterminedThe repository declares no quantization method, which normally means unquantized weights.
Trained context
UndeterminedThe config summary omits max_position_embeddings. Trained context is a model-card claim, not a derivable fact.
Declared licence
UndeterminedThe repository declares no license in its card metadata or tags. Absence of a declared license is not permission to use the weights.
Base model
UndeterminedThe repository declares no base model.
Library
transformerslibrary_name
Files in repository
12siblings
Last modified
2026-05-07lastModified

LICENCE POSTURE

Undeclared

The repository declares no licence. Absence of a licence is not permission: without a grant, default copyright applies to the weights.

This is a reading of a metadata field, not legal advice, and it does not account for the licences of upstream models or training data.

WEIGHT RESIDENCY FLOOR

The count below which it cannot fit.

Ceiling of checkpoint bytes over advertised accelerator memory. This is a lower bound on accelerator count for weights alone — KV cache, activations and runtime overhead all sit on top, so a real deployment needs more.

1×H100 80GB SXMNVIDIA
1×H200 141GB SXMNVIDIA
1×B200 180GBNVIDIA
1×A100 80GBNVIDIA
1×L40S 48GBNVIDIA
1×RTX 6000 Ada 48GBNVIDIA
1×GeForce RTX 5090 32GBNVIDIA
1×GeForce RTX 4090 24GBNVIDIA
1×Instinct MI300X 192GBAMD
1×Instinct MI325X 256GBAMD
1×Mac Studio M3 Ultra 512GBApple
1×MacBook Pro M4 Max 128GBApple

PROVENANCE

Derived from https://huggingface.co/api/models/trl-internal-testing/tiny-Qwen3MoeForCausalLM in the snapshot generated 2026-09-05. Evidence class hub_derived: computed mechanically from publisher metadata, never measured by this project.