Views
No views yet
mlc_llm package toolchain.use_function_calling: true)q4f16_0 layout (KN) is dramatically faster on Qualcomm Adreno GPUs compared to q4f16_1 (NK). The q4f16_1 format requires weight transposition operations that are catastrophically slow on Adreno's OpenCL implementation, resulting in 6-10x slower inference.mlc_llm package with mobile overrides (context_window_size=4096, prefill_chunk_size=1024, max_batch_size=1)model_lib: qwen2_q4f16_0_1be22ffdc6429c5019af9af8dae22086 when loading.