Views
No views yet
BF16 dtype-repack ofXiaomiMiMo/MiMo-V2.5-ASR-8B— original FP32 floating-point weights losslessly cast tobfloat16for LoRA / DoRA / PEFT compatibility and reduced disk footprint. The model architecture, parameter values, tokenizer, and configuration are identical to upstream — only the IEEE-754 storage dtype was changed.
License preserved end-to-end — seeLICENSEin this repo for the full text and attribution chain.
| Aspect | Upstream | This bundle |
|---|---|---|
| Floating-point storage dtype | FP32 | bfloat16 |
config.json torch_dtype | as-is | bfloat16 |
model.safetensors.index.json total_size | as-is | recomputed |
| Tokenizer / chat template / modeling code | as-is | unchanged |
| Number of parameters | as-is | unchanged |
| Value-level transformation beyond dtype cast | — | none |
| Disk size | 30 GB | 15 GB |
| Property | Value |
|---|---|
| Immediate parent | XiaomiMiMo/MiMo-V2.5-ASR-8B |
| Architecture | MiMoV2ASRForCausalLM |
| Architecture base / lineage | Qwen2-derived (model_type=qwen2) |
| Parameters | ~8B |
| Hidden size | 4096 |
| Num hidden layers | 36 |
| Attention heads / KV heads | 32 / 8 (GQA) |
| Vocab size | 151680 |
| Max position embeddings | 8192 |
| Format | bfloat16 |
| Bundle size on disk | 15 GB |
| License | MIT License |
| Project page | https://github.com/XiaomiMiMo/MiMo-V2.5-ASR |
repack_fp32_to_bf16.py
— reads each shard with safetensors.safe_open, casts floating-point
tensors to torch.bfloat16, rewrites the shard, updates the index
manifest. No GPU involvement, no value-level transformation
beyond the IEEE-754 dtype cast.dtype=torch.bfloat16 base1import torch
2from transformers import AutoTokenizer, AutoModelForCausalLM
3
4repo = "AMAImedia/MiMo-V2.5-ASR-8B-NOESIS-BF16"
5
6tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
7model = AutoModelForCausalLM.from_pretrained(
8 repo,
9 dtype=torch.bfloat16,
10 device_map="auto",
11 trust_remote_code=True,
12).eval()R-DTYPE-REPACK-BF16 — pure IEEE-754 dtype cast from FP32 to
bfloat16. No value-level transformation, no LoRA merge, no architectural
change. Equivalent to loading upstream with dtype=torch.bfloat16 and
saving, but materialised on disk.R-MIT-CLEAN — upstream MIT License preserved end-to-end via
the LICENSE file in this repo. AMAImedia adds only a derivative-work
notice for the repack step.R-NO-VALUE-TRANSFORM — no fine-tuning, no distillation, no merge has
been applied between upstream and this repo. Outputs are bit-for-bit
equivalent up to the precision difference of the dtype cast.XiaomiMiMo/MiMo-V2.5-ASR-8B. Original
model card, citation, and attribution from upstream apply without
modification. See LICENSE in this repo for the complete text plus the
NOESIS derivative-work NOTICE.1@misc{noesis2026mimov25asr8bnoesisbf16bf16,
2 title = {NOESIS DHCF-FNO :: MiMo-V2.5-ASR-8B-NOESIS-BF16 — BF16 dtype-repack derivative},
3 author = {Bolotnikov, Ilia and AMAImedia},
4 year = {2026},
5 note = {BF16 dtype-repack of XiaomiMiMo/MiMo-V2.5-ASR-8B for LoRA / PEFT
6 compatibility. 15 GB on disk, MIT License
7 preserved end-to-end.},
8 url = {https://huggingface.co/AMAImedia/MiMo-V2.5-ASR-8B-NOESIS-BF16}
9}LICENSE in this repo for citation requirements.AMAImedia/MiMo-V2.5-ASR-8B-NOESIS-BF16XiaomiMiMo/MiMo-V2.5-ASR-8B