The dataset econo-pairs-v2 is the second version of "econo-pairs". This dataset is built on top of the "samchain/bis_central_bank_speeches" dataset.
It is made of pairs of sequences of 512 tokens. The pairs are built following this approach:
If both sequences are coming from the same speech, the label is 1.0
If the sequences are coming from two different speeches but from the same central bank, the label is 0.5
If the sequences are coming from two different speeches… See the full description on the dataset page:
https://huggingface.co/datasets/samchain/econo-pairs-v2.