Views
No views yet
NemotronH_Nano_Omni_Reasoning_V3 — a tri-modal (text + vision + audio)
multimodal model:mlp1
projector).sound_projection).mlx_lm / mlx_vlm convert tools do not yet support the
NemotronH_Nano_Omni_Reasoning_V3 multimodal wrapper. The LLM-backbone tensor
layout is byte-identical to the text sibling
mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit;
the vision (vision_model.*) and audio (sound_encoder.*) towers are kept in
bf16. ~19 GB on disk, ~20 GB working set (Apple-silicon Mac tier).nemotron_h tooling. The vision and
audio towers require a multimodal runtime that implements the C-RADIO ViT-H and
Parakeet Conformer forward passes (e.g. the Evorix on-device engine).nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 under the
NVIDIA Open Model License; original model © NVIDIA.