Views
No views yet
Package from local in the Sinequa Runnable Model Wizard.| GPU | Quantization type | Batch size 1 | Batch size 32 |
|---|---|---|---|
| NVIDIA A10 | FP16 | 12 ms | 37 ms |
| NVIDIA T4 | FP16 | 20 ms | 71 ms |
| Quantization type | Memory |
|---|---|
| FP16 | 2300 MiB |
{
"inputs": [
{
"text": "What is the capital of China?"
},
{
"text": "Explain Gravity"
}
],
"options": {
"context": "query",
"passagePrefix": "",
"queryPrefix": "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:"
}
}max_length, under truncation to 1024.