CATALOGUE RECORD / HUB-DERIVED
MiniMaxAI
MiniMax-M2.7
Revision d494266a4affc0d2995ba1fa35c8481cbd84294b
TENSOR ACCOUNTING
Where the bytes are.
Summed from the safetensors index, one row per dtype. A parameter count alone cannot produce this figure, because a checkpoint may mix widths.
| Dtype | Parameters | Bytes each | Bytes | Share |
|---|---|---|---|---|
F8_E4M3 | 227,410,968,576 | 1 | 227.4 GB | 98.8% |
BF16 | 1,230,021,632 | 2 | 2.5 GB | 1.1% |
F32 | 48,774,656 | 4 | 195 MB | 0.1% |
EVERY FIELD, WITH ITS ORIGIN
Sourced or undetermined. Never assumed.
Each value below names the exact API field it was computed from. Where the Hub does not establish a value, the reason is shown instead of a plausible default.
- Architecture
- minimax_m2
config.model_type - Model classes
- MiniMaxM2ForCausalLM
config.architectures - Routed experts
- UndeterminedThe published config exposes no routed-expert count. The Hub's config summary omits fields some architectures place only in the full config.json.
- Experts per token
- 8
config.num_experts_per_tok - Shared experts
- UndeterminedThe published config declares no always-on shared experts.
- Routing sparsity
- UndeterminedRouting sparsity requires both a routed-expert and a per-token expert count.
- Total parameters
- 228,689,764,864
safetensors.total - Checkpoint bytes
- 230,066,110,464 (230.1 GB)
safetensors.parameters - Ships below 16-bit
- Yes
safetensors.parameters - Quantisation method
- fp8
config.quantization_config.quant_method - Trained context
- UndeterminedThe config summary omits max_position_embeddings. Trained context is a model-card claim, not a derivable fact.
- Declared licence
- other
cardData.license - Base model
- UndeterminedThe repository declares no base model.
- Library
- transformers
library_name - Files in repository
- 151
siblings - Last modified
- 2026-04-20
lastModified
LICENCE POSTURE
Vendor terms
The repository ships bespoke vendor terms rather than a standard open licence. Acceptable-use clauses, attribution duties and user-count thresholds are common; read the terms in full.
This is a reading of a metadata field, not legal advice, and it does not account for the licences of upstream models or training data.
WEIGHT RESIDENCY FLOOR
The count below which it cannot fit.
Ceiling of checkpoint bytes over advertised accelerator memory. This is a lower bound on accelerator count for weights alone — KV cache, activations and runtime overhead all sit on top, so a real deployment needs more.