Views
No views yet
onnx/encoder_model{_fp16,_q4}.onnx — mel encoder (input_features [1, 128, T] → last_hidden_state)onnx/decoder_model_merged{_fp16,_q4}.onnx — KV-cache merged decoder (use_cache_branch)
with cross_attentions.{i} outputs for DTW word timestamps against the model's
supervised alignment heads (generation_config.json)<|startoftranscript|>):[verbatim_1][verbatim_2][verbatim_3][verbatim_4][verbatim_5] (<htx> hotwords <ehtx>)? <|startoftranscript|> <|lang|> <|transcribe|> <|notimestamps|>[UM], [UH], [laughter],
[cough], ...). Intended mode uses [intended_1..5] tags instead. See the
upstream model card and
nyrahealth/CrisperWhisper for details.nyralabs/CrisperWhisper2.0_turbo and are
distributed under the Nyra Health Non-Commercial Research License (see LICENSE.md
in this repository — identical to the upstream license). In short:LICENSE.md.main_export (task automatic-speech-recognition-with-past,
output_attentions=True), fp16 via onnxconverter-common (keep_io_types=True),
q4 via onnxruntime MatMulNBitsQuantizer(bits=4, block_size=32, is_symmetric=True).