Views
No views yet
USER: <image>\nIs at least one {noun} visible in this image? Answer yes or no.\nASSISTANT:| detector (same eval universe) | AP | F1 | within-word AUROC |
|---|---|---|---|
| LLaVA-13B self logit (same forward pass) | .740 | — | — |
| 122B VLM judge (logit readout) | .826 | .777 | .930 |
| verify probe 7B (3-seed ens) | .876 [.862, .889] | .814 | .937 |
| verify probe 13B (this repo, 3-seed ens) | .883 [.870, .895] | .817 | .940 |
probe_verify_contrast_13b_s{0,1,2}.pt).probe_verify_contrast_13b_s0.pt, ..._s1.pt, ..._s2.pt — PyTorch
state dicts of the per-layer MLP heads (one PairMLP each).probe_config.json — layers, dims, feature mode, metrics.1import torch, torch.nn as nn
2
3LAYERS = [10, 15, 20, 25, 30]
4
5class PairMLP(nn.Module):
6 def __init__(self, d_in=5120, hidden=256):
7 super().__init__()
8 self.mlp = nn.ModuleList([
9 nn.Sequential(nn.Linear(d_in, hidden), nn.GELU(), nn.LayerNorm(hidden),
10 nn.Linear(hidden, hidden), nn.GELU(), nn.LayerNorm(hidden),
11 nn.Linear(hidden, 1))
12 for _ in LAYERS])
13
14 def forward(self, X): # X: {layer: (N, d_in) contrast features}
15 return [m(X[l]).squeeze(-1) for m, l in zip(self.mlp, LAYERS)]
16
17models = []
18for s in (0, 1, 2):
19 m = PairMLP()
20 m.load_state_dict(torch.load(f"probe_verify_contrast_13b_s{s}.pt",
21 map_location="cpu"))
22 m.eval()
23 models.append(m)
24
25# p_halluc = mean over seeds of (mean over layers of sigmoid(head(x)))
26with torch.no_grad():
27 p = torch.stack([
28 torch.stack([torch.sigmoid(z) for z in m(X)]).mean(0)
29 for m in models]).mean(0)general_hallucination/scripts/cocogt/verify_extract.py in the training repo
(run with --base_model llava-hf/llava-1.5-13b-hf --layers 10,15,20,25,30);
fitting/eval is verify_fit.py in the same directory.