InspectAI evaluation logs from the Benchmarking aLLarMa study, which evaluates 9 open-weight LLMs (0.8B–20B parameters) on a two-stage constrained GraphRAG framework. Every run was repeated on two reference machines so the results can be compared between a consumer GPU workstation and a low-power edge appliance.
Stage 1 — Retriever. 81 retrieval strategies — 58 non-LLM baselines, 21 LLM-augmented, and 2… See the full description on the dataset page:
https://huggingface.co/datasets/imsaumil/allarma-benchmark.