nemotron-3-embed-8b-legal
A specialized embedding model for the legal domain, fine-tuned from
nvidia/Nemotron-3-Embed-8B for judgments, contracts, statutes and regulatory holdings.
It scores
84.62 mean NDCG@10 across all 12 datasets of the
Massive Legal Embedding
Benchmark (MLEB), which is
first on the benchmark, ahead of
Kanon 2 Embedder at 81.86. Evaluated with the benchmark's official harness at its pinned
commit; results are
submitted to the benchmark
maintainers and listing is pending their review.
On
MTEB(Law, v1) it ranks
1st by Rank (Borda), the leaderboard's default ranking, among the 158 models with
complete coverage of the benchmark, and 2nd by Mean (Task) at 69.33. Per task it is 1st of
162 on LegalBench Consumer Contracts QA, 2nd of 200 on Corporate Lobbying, 2nd of 156 on
GerDaLIRSmall and 4th of 210 on Legal Summarization (standings as of 27 July 2026).
Model Details
| |
|---|
| Base model | nvidia/Nemotron-3-Embed-8B-BF16 |
| Parameters | 8B |
| Embedding dimension | 4,096 |
| Max sequence length | 32,768 tokens |
| Pooling / similarity | mean pooling, cosine |
| Language | English |
| License | OpenMDW-1.1 (same as base) |
Usage
1from sentence_transformers import SentenceTransformer
2
3model = SentenceTransformer("minetta/nemotron-3-embed-8b-legal")
4
5queries = model.encode(
6 ["did the court find a violation of Article 3?"],
7 prompt_name="query",
8)
9documents = model.encode(
10 ["FOR THESE REASONS, THE COURT unanimously..."],
11 prompt_name="document",
12)
13
14scores = queries @ documents.T
The query: / passage: prompts are built into the model config; use them as shown. Works
well for retrieval over case law, contract clauses, statutes and regulatory decisions, at
paragraph-to-document granularity.
Performance
MLEB — Massive Legal Embedding Benchmark
Mean NDCG@10 across all 12 datasets. Published leaderboard scores shown for reference
(as of July 2026); our results are submitted to the maintainers for review.
| model | score |
|---|
| nemotron-3-embed-8b-legal (this model, self-reported) | 84.62 |
| Kanon 2 Embedder | 81.86 |
| voyage-4-large | 81.05 |
| Nemotron-3-Embed-8B (base) | 80.62 |
MLEB per-dataset results (as of July 2026)
| dataset | base | this model |
|---|
| echr-retrieval | 43.13 | 83.37 |
| scalr | 76.21 | 79.90 |
| contractual-clause-retrieval | 88.36 | 91.27 |
| legal-rag-bench | 64.91 | 67.42 |
| australian-tax-guidance-retrieval | 80.75 | 82.35 |
| license-tldr-retrieval | 79.43 | 80.90 |
| bar-exam-qa | 63.61 | 64.29 |
| irish-legislative-summaries | 94.19 | 94.33 |
| consumer-contracts-qa | 90.80 | 90.78 |
| singaporean-judicial-keywords | 96.75 | 96.34 |
| uk-legislative-long-titles | 97.20 | 96.40 |
| gdpr-holdings-retrieval | 92.10 | 88.06 |
MTEB — legal retrieval
Evaluated with mteb main@787bf7c via mteb.get_model(...) — the reference implementation
in the official registry. Standings are against every model in the MTEB results repository
carrying a score on that task, as of 26 July 2026.
| task | score | standing |
|---|
| LegalBenchConsumerContractsQA | 87.81 | 1st of 162 |
| LegalBenchCorporateLobbying | 96.27 | 2nd of 200 |
| GerDaLIRSmall | 48.59 | 2nd of 156 |
| LegalSummarization | 77.16 | 4th of 210 |
| LegalQuAD | 73.17 | 5th of 202 |
| AILACasedocs | 49.13 | 10th of 206 |
| AILAStatutes | 58.65 | 18th of 232 |
| LeCaRDv2 | 63.81 | 28th of 157 |
| BarExamQA | 65.36 | 1st of 3 — new task, few models evaluated |
| BillSumCA | 99.02 | 1st of 3 — new task, few models evaluated |
| BillSumUS | 99.93 | 1st of 3 — new task, few models evaluated |
MTEB(Law, v1): 1st by Rank (Borda), the leaderboard's default ranking, and 2nd by
Mean (Task) at 69.33, among the 158 models with complete coverage of the benchmark, as
of 27 July 2026. All eight member tasks were evaluated. Results in
embeddings-benchmark/results#640.
MLEB scores for this model (and the base model) are our local runs of the official MLEB
harness at its pinned commit; other models' MLEB numbers are their published leaderboard
scores (July 2026). Data handling and the train/eval separation audit are documented in
CONTAMINATION.md and
PROVENANCE.md.
Training
Fine-tuned with LoRA stages and a contrastive objective (hard-negative mining) on public
legal sources: ECtHR case law, GDPR enforcement decisions, SEC EDGAR agreements, US court
opinions, Australian legal corpora, and terms-of-service texts. Benchmark evaluation data
was excluded from training and the exclusion verified per-document. Training pair files are
not redistributed; source and license details are in
PROVENANCE.md.
Limitations
- One MLEB dataset (gdpr-holdings-retrieval) scores below the base model (see the
per-dataset table above).
- English only; strongest in the jurisdictions covered by the training sources (EU, US, UK,
Australia).
- An embedding model for retrieval, not legal advice.
License
OpenMDW-1.1, inherited from the base model; the base's LICENSE and notices ship in this
repository. Community fine-tune; not affiliated with NVIDIA.