Views
No views yet
| Eval | Shisa V1 7B V2.1 | Shisa 7B V1 | Shisa Gamma 7B V1 | Shisa V2 8B | Shisa V2.1 8B |
|---|---|---|---|---|---|
| JA AVG | 41.5 | 26.2 | 38.0 | 58.7 | 67.8 |
| EN AVG | 28.5 | 29.4 | 21.1 | 55.1 | 57.8 |
| Shaberi v2.1 | 5.223 | 3.743 | 5.511 | 6.427 | 7.353 |
| ELYZA 100 | 5.590 | 4.050 | 5.723 | 7.300 | 7.660 |
| JA MT-Bench | 4.892 | 3.326 | 4.989 | 5.975 | 7.783 |
| Rakuda | 6.075 | 3.123 | 5.947 | 6.463 | 7.150 |
| Tengu | 4.335 | 4.473 | 5.386 | 5.970 | 6.817 |
| M-IFEval (JA) | 0.343 | 0.151 | 0.256 | 0.477 | 0.471 |
| shisa-jp-ifeval | 0.133 | 0.093 | 0.107 | 0.293 | 0.347 |
| shisa-rp-bench | 2.159 | 1.547 | 3.225 | 4.739 | 4.792 |
| shisa-tl-bench | 4.825 | 0.664 | 0.575 | 7.617 | 8.917 |
| kiseki-eval | 2.359 | 2.439 | 2.447 | 3.348 | 3.580 |
| chotto-eval | 0.200 | 0.018 | 0.036 | 0.145 | 0.455 |
| MixEval Easy | 0.521 | 0.639 | 0.503 | 0.821 | 0.802 |
| MixEval Hard | 0.340 | 0.361 | 0.217 | 0.555 | 0.607 |
| LiveBench | 10.8 | 14.2 | 11.8 | 32.1 | 45.7 |
| GPQA Diamond | 0.086 | 0.121 | 0.086 | 0.379 | 0.328 |
| IFEval | 0.335 | 0.340 | 0.274 | 0.828 | 0.791 |
| IFBench | 0.167 | 0.170 | 0.143 | 0.330 | 0.259 |
| HumanEval+ | 0.439 | 0.287 | 0.134 | 0.622 | 0.805 |
gpt-5.1-2025-11-13).
loose score.mixeval_easy task in Lighteval/Inspect) combining free‑form and multiple‑choice questions with 0.96 correlation with 2024 Chatbot Arena rankings; scored both by the task’s exact metrics and by LLM judges (Flow‑Judge flowaicom/Flow-Judge-v0.1 and a GPT judge, default gpt-4.1-mini-2025-04-14) via the HF Lighteval runner.mixeval_hard) designed to better separate strong models, run through the same Lighteval/Inspect pipeline and Flow‑Judge + GPT‑judge scoring as MixEval Easy.LiveBench-2024-11-25. Our fork supports concurrent runs, GPT-5.1 reasoning semantics, and other fixes.gpqa:diamond task using Idavidrein/gpqa); we score with Inspect’s multiple‑choice choice metric and an additional robustness pass that recovers bare letter answers, so this remains a pure reference‑based metric (no LLM judge).ifeval task in Lighteval/Inspect over google/IFEval), scored with the original rule‑based check_following functions; we report the loose prompt‑level accuracy, with no LLM judge involved.loose prompt-level accuracies using IFBench's own verification functions (no LLM judge). We fix some evaluation bugs and add a response generation script for parallel execution against OpenAI-compatible endpoints.plus-pass@1 score (reference/test-based judgement).