nvidia/parakeet-tdt-0.6b-v3.litert-torch commit
4502654d6afe38b495d2efe8dcd37e057d0c5e7f and the
dynamic_wi8_afp32 quantization recipe.encode_quantized.tflite: [1, 128, 1001] → [1, 1024, 126]decode_quantized.tflite: encoder output, one float32 token ID, and two
[2, 1, 640] states → [1, 126, 1, 8198] logits and next statesCompiledModel API can use uniform float buffers. Split files
allow independent NPU/CPU placement.