⚠️ Measured on a slice of the dataset, not the full suite: this score is not comparable with a full-suite run.
⚡ Measured Speed on the Publishing Machine
Measured on NVIDIA GeForce RTX 5070 Ti • 15.9 GB VRAM. Two different numbers follow, and they are not interchangeable.
What was measured
Value
How
Single-stream decode (what a chat feels)
200.2 tok/s
one request at a time, on NVIDIA GeForce RTX 5070 Ti
Prompt processing
1904 tok/s
same probe
Aggregate throughput during evaluation
192.1 tok/s
several requests in flight — not what a single answer runs at
Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.
🚀 Quick Start Guide
1. Running with Sigma Studio (Recommended)
Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:
bash
1# Clone and run Sigma Studio2git clone https://github.com/Sigmanih/SigmaStudio.git
3cd SigmaStudio
4.\sigma_studio.bat
2. Running with llama.cpp
llama-cli -hf sigmanih/sigma-alpaca-3b-gguf -p "Hello! How can I help you today?" -ngl 99
🇮🇹 Documentazione in Italiano
sigma-alpaca-3b-gguf è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso Σ-SIGMA Studio.
📋 Specifiche e Configurazione
Architettura Base:llama (3B parametri)
Formato Pesi:GGUF (Q4_K_M)
Spazio su Disco:1.88 GB
Finestra di Contesto:131,072 token
Profilo d'Uso Consigliato: Dispositivi edge, agenti vocali in tempo reale, CPU e carichi leggeri.
⚠️ Misurato su una porzione del dataset, non sulla suite intera: il punteggio non è confrontabile con uno ottenuto sull'intero.
Data Test:2026-08-30 su motore deterministico SigmaEngine
⏱️ Throughput Hardware e Fasce Consigliate
Velocità Verificata in Locale:192.1 tok/s su NVIDIA GeForce RTX 5070 Ti.
Risposta singola (quello che si sente in chat):200.2 tok/s su NVIDIA GeForce RTX 5070 Ti.
Lettura del prompt:1904 tok/s.
Throughput complessivo durante la valutazione:192.1 tok/s — piu' richieste in volo insieme, non la velocita' di una risposta singola.
Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.
⭐ Supporta il Progetto Open Source
Se questo modello ti è utile o vuoi esplorare l'ecosistema completo: