Views
No views yet
kyutai/moshiko-pytorch-bf16@2bfc9ae6e89079a5cc7ed2a68436010d91a3d289onnxruntime/mobius@80e11c45cb47b876bc2b4c6b183ebf38427c121b.provenance.json.inference_metadata.yaml.graph_report.json.1hf download justinchuby/moshiko-full-duplex-onnx-catalogue \
2 --repo-type model --local-dir moshiko-full-duplex
3cd moshiko-full-duplex1pip install 'onnxruntime-gpu==1.28.0' 'numpy==2.3.3' 'soundfile==0.13.1' \
2 'nvidia-cudnn-cu12==9.10.2.21' 'nvidia-cublas-cu12==12.6.4.1' \
3 'nvidia-cuda-runtime-cu12==12.6.77'
4bash run_cuda.sh --seed 0 --dep-q 8 --frames 5 --save-to reproducedoutput.npz contains generated text tokens, 8-stream audio codes, and decoded waveform; assistant.wav is the decoded sample. On H200, compact INT4 steady-state LM frame time averaged 19.97 ms versus the 80 ms budget. Exact per-frame timings are in runtime_output.json.inference_metadata.annotated.yaml for inline explanations of this package's workflow, tensor/state/cache contracts, and fail-closed omissions. inference_metadata.yaml remains the canonical machine-authored contract; automated validation confirms both files parse to the same metadata object.