A preprocessed retrieval corpus built from the International Military Tribunal (IMT) proceedings at Nuremberg (1945–1946), structured for hybrid dense-sparse RAG pipelines. 46,325 chunks covering daily trial transcripts, prosecution documents, the final judgment, and supporting briefs.
This corpus underpins Nuremberg Scholar, a RAG system using BGE-M3 hybrid retrieval, bge-reranker-v2-m3 cross-encoder reranking, and Llama-3.1-8B-Instruct generation.… See the full description on the dataset page:
https://huggingface.co/datasets/dtufail/nuremberg-trials-corpus.