Views
No views yet
[numTokens, 128]). Relevance is computed with MaxSim (for each query
token, the max dot product over document tokens, summed). Use it when you want
stronger retrieval quality than single-vector embeddings — e.g. RAG / search..pte files; with
React Native ExecuTorch the library constants guarantee this."[Q] " for queries,
"[D] " for documents.forward to get the per-token [S, 128] matrix for each text.config_sentence_transformers.json) tokenizes to:
[510..524, 535..541, 568..573, 600..603] (32 ids).xnnpack/, mlx/ — the partitioned .pte files + per-backend config.json.tokenizer.json — wire to tokenizerSource.config.json, tokenizer_config.json — reference metadata.