This dataset is similar to Free-Law-Project/opinions-synthetic-query-512, the only difference is the opinions are chunked to at most 7800 tokens instead of 480 tokens, tokenized using the bert-base-cased tokenizer with 2 sentence overlap. The number of tokens is just shy of the 8192 context window limit to account for tokenization variation between the different encoder models for experiments.
The dataset is used to finetune the semantic search model… See the full description on the dataset page:
https://huggingface.co/datasets/freelawproject/opinions-synthetic-query-8192.