tokenfold-select
A LoRA adapter for ibm-granite/granite-embedding-reranker-english-r2.
It ranks text spans by their relevance to a query, encoding each (query, span) pair jointly and
returning one logit. Higher logits indicate stronger evidence that the span should be kept.
This is a ranking model, not a standalone compressor. Use it with an allocator that separately
preserves required content and enforces the token budget. Do not use its scores as a safety filter.
Intended use
Use the model to rank pre-segmented passages, log lines, code blocks, or other text spans before
assembling a smaller context. Scores are most useful for ordering spans from the same document;
they are not calibrated probabilities or a substitute for hard retention rules.
How to use
1from pathlib import Path
2
3import torch
4from huggingface_hub import snapshot_download
5from transformers import AutoModelForSequenceClassification, AutoTokenizer
6from peft import PeftModel
7
8base_id = "ibm-granite/granite-embedding-reranker-english-r2"
9repo_dir = Path(snapshot_download("OWNER/tokenfold-select")) # replace with this repository ID
10adapter_dir = repo_dir / "adapter"
11tok = AutoTokenizer.from_pretrained(adapter_dir)
12base = AutoModelForSequenceClassification.from_pretrained(base_id, dtype=torch.float32)
13model = PeftModel.from_pretrained(base, adapter_dir).eval()
14
15def score(query: str, spans: list[str]) -> list[float]:
16 if not spans:
17 return []
18 enc = tok([query] * len(spans), spans, padding=True, truncation=True,
19 max_length=8192, return_tensors="pt")
20 with torch.no_grad():
21 out = model(input_ids=enc["input_ids"], attention_mask=enc["attention_mask"])
22 return out.logits.view(-1).float().tolist()
The repository contains a PEFT adapter, not the base-model weights. Loading therefore also
downloads ibm-granite/granite-embedding-reranker-english-r2.
Training data
97,449 source/query fixtures spanning code, logs, diffs, JSON and tool calls, agentic tool
use, and long-context QA. Sources include project-authored synthetic examples, SWE-bench Verified,
publicly available tool-output benchmark samples, and samples derived from HotpotQA, NarrativeQA,
SQuAD, TriviaQA, and MS MARCO. Each fixture pairs a source document, a query, a gold answer span,
and optional required spans that the downstream allocator must preserve regardless of model score.
Training procedure
LoRA (r=8, alpha=16, dropout=0.05, target_modules="all-linear") via peft, applied for two epochs
to ibm-granite/granite-embedding-reranker-english-r2 with BCEWithLogitsLoss and class-weighted positives. The released adapter was
trained on the full corpus; the results below come from separately trained held-out evaluation
runs.
Evaluation
Mean task success over three stratified repeated-subsampling runs (about 73,000 training and 24,000
held-out fixtures per run). Each run fine-tuned a fresh adapter, evaluated every method on the same
held-out fixtures, and used the same required-span and token-budget allocator. Task success means
that the literal gold-answer span survived compression.
| target token ratio | this model | Kompress-v2 (native) | Kompress-v2 relevance scorer | BM25 |
|---|
| 0.5 | 0.863 | 0.665 | 0.805 | 0.794 |
| 0.25 | 0.703 | 0.472 | 0.615 | 0.607 |
| 0.1 | 0.399 | 0.304 | 0.377 | 0.377 |
At matched forced budgets, this model outperformed each baseline listed above at all three ratios.
Kompress-v2 and BM25 are named here as benchmark baselines the model is compared against, not as
an influence on this model's design.
Limitations
- Training labels are weak per-unit signals (does this span contain the gold answer?), not judged
per-token labels — treat results as directional.
- Task success measures literal answer-span retention, not downstream answer quality.
- Fixtures are English-centric despite the multilingual base model.
- Inputs longer than 8,192 tokens are truncated.
- Logits are uncalibrated and should be used for ranking, not as probabilities.
- The model cannot guarantee preservation of required or safety-critical text; the downstream
allocator must enforce those guarantees.
License
The adapter is released under Apache 2.0. The base model is also Apache 2.0; source datasets remain
subject to their own terms.