Views
No views yet
| Backend | Quantization | Max sequence length | Init time (ms) | Inference time (ms) | Memory (RSS in MB) | Model size (MB) |
|---|---|---|---|---|---|---|
NPU | Mixed Precision* | 256 | 206 | 7.8 | 224 | 182 |
NPU | Mixed Precision* | 512 | 241 | 18 | 231 | 184 |
NPU | Mixed Precision* | 1024 | 272 | 57 | 263 | 195 |
NPU | Mixed Precision* | 2048 | 468 | 169 | 332 | 220 |
GPU | Mixed Precision* | 256 | 1175 | 64 | 762 | 179 |
GPU | Mixed Precision* | 512 | 1445 | 119 | 762 | 179 |
GPU | Mixed Precision* | 1024 | 1545 | 241 | 771 | 183 |
GPU | Mixed Precision* | 2048 | 1707 | 683 | 786 | 196 |
CPU | Mixed Precision* | 256 | 17.6 | 66 | 110 | 179 |
CPU | Mixed Precision* | 512 | 24.9 | 169 | 123 | 179 |
CPU | Mixed Precision* | 1024 | 35.4 | 549 | 169 | 183 |
CPU | Mixed Precision* | 2048 | 35.8 | 2455 | 333 | 196 |
Install Google Tensor SDK section within the LiteRT_AOT_Compilation_Tutorial.ipynb notebook.