Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
gentoo-small-model-throughput-vllm – Dataset by LostGentoo | AlphaNeural AI
You can deploy this model and start earning money today!
LostGentoo
/
gentoo-small-model-throughput-vllm
like
0
text-generation
en
mit
us
benchmark
throughput
quality
lighteval
vllm
qwen3.5
lfm2
minicpm
Views
No views yet
Model card
Files and Versions
Community
API
Small-model throughput + quality benchmarks (RTX 5060 Ti 16 GB)
Local vLLM throughput/scaling measurements and lighteval accuracy benchmarks for sub-2B chat models on a 3× RTX 5060 Ti (16 GB) host. Raw JSON outputs plus comparison tables.
Setup
Setting Value
Host RTX 5060 Ti 16 GB (CUDA device 2)
GPU CUDA device 2
Engine vLLM 0.24
max_model_len 4096
gpu_memory_utilization 0.75
max_num_seqs 128
Gen length 128 tokens
Warmup / repeats 2 / 2… See the full description on the dataset page:
https://huggingface.co/datasets/LostGentoo/gentoo-small-model-throughput-vllm
.