Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3-Embedding-0.6B-vllm-W8A8 – AI Model by dolfsai | AlphaNeural AI
You can deploy this model and start earning money today!
dolfsai
/
Qwen3-Embedding-0.6B-vllm-W8A8
like
0
safetensors
qwen3
feature-extraction
en
es
HuggingFaceH4/ultrachat_200k
Qwen/Qwen3-Embedding-0.6B
quantized
apache-2.0
8-bit
compressed-tensors
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
prudant/Qwen3-Embedding-0.6B-W8A8
This is a compressed version of Qwen/Qwen3-Embedding-0.6B using llm-compressor with the following scheme: W8A8
Important
: You MUST read the following guide for correct usage of this model here
Guide
Model Details
Original Model
: Qwen/Qwen3-Embedding-0.6B
Quantization Method
: GPTQ
Compression Libraries
:
llm-compressor
Calibration Dataset
: ultrachat_200k (1024 samples)
Optimized For
: Inference with vLLM
License
: same as original model