Views
No views yet
deepseek-ai/DeepSeek-V2-Lite-Chat (frontier-panel, BOUNDARY regime)deepseek-ai/DeepSeek-V2-Lite-Chat trunk. Configuration: mid_dim=128, K=4 (frontier-panel, BOUNDARY regime).Calibrate the sign before deploying. The direction of the disagreement signal is trunk-family-specific: on some families it rises on out-of-distribution input, on others (Llama-70B-class, gpt-oss-120B) it falls. Score ~50 known-ID and ~50 known-OOD prompts once and check which direction separates. Details: the paper and repo below.
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3from eu_halt import attach
4
5model = AutoModelForCausalLM.from_pretrained(
6 "deepseek-ai/DeepSeek-V2-Lite-Chat", torch_dtype=torch.bfloat16,
7).to("cuda").eval()
8tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V2-Lite-Chat")
9
10uncertainty = attach(
11 model,
12 heads_repo="debajyotidasgupta/eu-halt-deepseek-v2-lite-chat",
13 mid_dim=128,
14)
15print(uncertainty("Who founded Quora in 2008?", tokenizer))
16# Higher = more uncertain.model.safetensors — the K=4 head weights (preferred format; config embedded as metadata).config.json — head geometry: num_heads, mid_dim, source_layers, base_model, dims.heads_final.pt — the original torch checkpoint (kept for backward compatibility).heads_step{500,1000,1500,2000,2500}.pt — intermediate checkpoints (where uploaded).source_layers.json — the K=4 trunk-layer indices the heads read from.history.json — per-step loss + disagreement + GPU stats.HuggingFaceFW/fineweb-edu (streaming).quiet variants).| Signal | AUROC | 95% CI |
|---|---|---|
| disagreement | 0.5451 | [0.4612, 0.6326] |
| entropy | 0.6012 | [0.5394, 0.6607] |
| last_token_unc | 0.1663 | [0.1076, 0.2308] |
| mahalanobis | 0.3160 | [0.1703, 0.4876] |
| targ_margin | 0.1685 | [0.1103, 0.2333] |
| etc_trend | 0.5078 | [0.4426, 0.5724] |
| llm_check | 0.1740 | [0.1104, 0.2564] |
| mc_dropout | 0.4392 | [0.3606, 0.5138] |
| rauq | nan | [nan, nan] |
| p_true | nan | [nan, nan] |
| semantic_entropy | nan | [nan, nan] |
| semantic_entropy_nli | nan | [nan, nan] |
| eigenscore | nan | [nan, nan] |
uncertainty.per_token(text, tokenizer).deepseek-ai/DeepSeek-V2-Lite-Chat retains its own license (Qwen3 / Llama-3 / Phi / Gemma).1@inproceedings{dasgupta2026euhalt,
2 author = {Dasgupta, Debajyoti and Mondal, Arijit and Chakrabarti, Partha P.},
3 title = {When Uncertainty Lies: How Model Scale and Layer Geometry Quietly
4 Invert the Meaning of Internal Disagreement in Large Language Models},
5 booktitle = {1st Conference For AI Scientists (CAISc)},
6 year = {2026},
7 url = {https://huggingface.co/debajyotidasgupta/eu-halt-deepseek-v2-lite-chat},
8}