LF_BERT_v1 is a lightweight TinyBERT-based cross-encoder fine-tuned for semantic evidence filtering in Retrieval-Augmented Generation (RAG) pipelines.
The model acts as a semantic gatekeeper, scoring (query, candidate_sentence) pairs to determine whether the sentence is factually useful evidence or a semantic distractor.
It is designed for CPU-only, edge, and offline deployments, with millisecond-level inference latency.
This model is the core filtering component of Project Sentinel.
0 – Distractor: Topically similar but factually insufficient sentences
Dataset Statistics
Split
Samples
Train
69,101
Validation
7,006
The dataset is intentionally imbalanced, reflecting real retrieval scenarios.
Training Procedure
Hyperparameters
Learning rate: 1e-5
Batch size: 16
Epochs: 2
Optimizer: AdamW
Scheduler: Linear
Seed: 42
Loss: Weighted cross-entropy
Training Results
Epoch
Validation Loss
F1
Accuracy
Precision
Recall
ROC-AUC
1
0.4003
0.7119
0.8290
0.6146
0.8457
0.9038
2
0.4042
0.7028
0.8167
0.5907
0.8674
0.9064
Thresholded Performance (Strict Guard)
Decision threshold: 0.85
Hallucination rate: 5.92%
Fact retention: 60.34%
Average latency: 5.30 ms (CPU)
This configuration prioritizes trustworthiness over recall.
Citation
If you use this model, please cite:
@article{salih2026sentinel,
title={Project Sentinel: Lightweight Semantic Filtering for Edge RAG},
author={Salih, El Mehdi and Ait El Mouden, Khaoula and Akchouch, Abdelhakim},
year={2026}
}