Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
tmmluplus-local-llm-results – Dataset by betty0 | AlphaNeural AI
You can deploy this model and start earning money today!
betty0
/
tmmluplus-local-llm-results
like
0
question-answering
zh
mit
1K<n<10K
json
tabular
text
datasets
dask
polars
mlcroissant
us
llm-evaluation
tmmluplus
ollama
benchmark
traditional-chinese
taiwan
zh-tw
Views
No views yet
Model card
Files and Versions
Community
API
TMMLU+ Local LLM Arena — 本機開源模型評測結果
在單張 RTX 4090 上以 Ollama 對 5 個開源模型(6 個受測 configuration,Qwen3 另分 thinking on/off 兩組)跑 ikala/tmmluplus (台灣版 MMLU)的 zero-shot 四選一評測結果。原始碼與完整報告:
https://github.com/tun0000/tmmluplus-local-arena
排行榜
模型 正確率 % 無效輸出 % 平均延遲 s
Qwen3 8B (think) 69.8 2.4 19.0
DeepSeek-R1 8B (think) 57.4 1.6 20.8
Qwen3 8B (no-think) 55.0 0.0 2.4
Gemma 3 12B 53.8 0.0 2.7
GLM-4 9B 48.2 0.0 2.4
Llama 3.1 8B 41.6 0.0 2.5
完整各科分數見… See the full description on the dataset page:
https://huggingface.co/datasets/betty0/tmmluplus-local-llm-results
.