CATALOGUE RECORD / HUB-DERIVED

nvidia

Qwen3.6-35B-A3B-NVFP4

Revision 1355db6a052410cfd62085d94b58866fd0f2c3c5

—TOTAL PARAMETERS
21.4 GBCHECKPOINT BYTES
— / —EXPERTS / PER TOKEN
10.3MDOWNLOADS

INDEX INCONSISTENCY

This repository publishes two different parameter counts.

The safetensors index declares 18,683,860,336 parameters in its total, while its own per-dtype map sums to 19,528,501,104. Those two fields describe the same tensors, so one of them is wrong.

No parameter count is published for this record. Picking the more plausible of two contradictory figures would be a guess presented as a fact. The tensor byte total below is computed from the per-dtype map alone and is cross-checked against the repository’s stored bytes.

TENSOR ACCOUNTING

Where the bytes are.

Summed from the safetensors index, one row per dtype. A parameter count alone cannot produce this figure, because a checkpoint may mix widths.

DtypeParametersBytes eachBytesShare
U816,423,321,600116.4 GB76.9%
BF161,825,916,78423.7 GB17.1%
F8_E4M31,279,262,72011.3 GB6.0%

EVERY FIELD, WITH ITS ORIGIN

Sourced or undetermined. Never assumed.

Each value below names the exact API field it was computed from. Where the Hub does not establish a value, the reason is shown instead of a plausible default.

Architecture
qwen3_5_moeconfig.model_type
Model classes
Qwen3_5MoeForConditionalGenerationconfig.architectures
Routed experts
UndeterminedThe published config exposes no routed-expert count. The Hub's config summary omits fields some architectures place only in the full config.json.
Experts per token
UndeterminedThe published config exposes no per-token expert count.
Shared experts
UndeterminedThe published config declares no always-on shared experts.
Routing sparsity
UndeterminedRouting sparsity requires both a routed-expert and a per-token expert count.
Total parameters
UndeterminedThe published index is internally inconsistent: safetensors.total declares 18,683,860,336 parameters while the per-dtype map sums to 19,528,501,104. Neither figure can be treated as the parameter count.
Checkpoint bytes
21,354,417,888 (21.4 GB)safetensors.parameters
Ships below 16-bit
Yessafetensors.parameters
Quantisation method
modeloptconfig.quantization_config.quant_method
Trained context
UndeterminedThe config summary omits max_position_embeddings. Trained context is a model-card claim, not a derivable fact.
Declared licence
apache-2.0cardData.license
Base model
Qwen/Qwen3.6-35B-A3BcardData.base_model
Library
Model Optimizerlibrary_name
Files in repository
17siblings
Last modified
2026-08-29lastModified

LICENCE POSTURE

Permissive

The repository declares a licence that is generally read as permitting commercial use. Read the licence file in the repository before relying on that.

This is a reading of a metadata field, not legal advice, and it does not account for the licences of upstream models or training data.

WEIGHT RESIDENCY FLOOR

The count below which it cannot fit.

Ceiling of checkpoint bytes over advertised accelerator memory. This is a lower bound on accelerator count for weights alone — KV cache, activations and runtime overhead all sit on top, so a real deployment needs more.

1×H100 80GB SXMNVIDIA
1×H200 141GB SXMNVIDIA
1×B200 180GBNVIDIA
1×A100 80GBNVIDIA
1×L40S 48GBNVIDIA
1×RTX 6000 Ada 48GBNVIDIA
1×GeForce RTX 5090 32GBNVIDIA
1×GeForce RTX 4090 24GBNVIDIA
1×Instinct MI300X 192GBAMD
1×Instinct MI325X 256GBAMD
1×Mac Studio M3 Ultra 512GBApple
1×MacBook Pro M4 Max 128GBApple

PROVENANCE

Derived from https://huggingface.co/api/models/nvidia/Qwen3.6-35B-A3B-NVFP4 in the snapshot generated 2026-09-05. Evidence class hub_derived: computed mechanically from publisher metadata, never measured by this project.