These conversations are extracted from Gutenberg books using code and LLM extraction.
They are guaranteed to be identical to the original text (every dialog text is matched against the original book letter by letter), but it's not guaranteed that all dialogs are extracted, or that the speaker is correct.
The nearby words of the same speaker are grouped together, eg:
“He has a thirst for travelling; perhaps he may turn out a Bruce or a
Mungo Park,” said Mr.… See the full description on the dataset page:
https://huggingface.co/datasets/croqaz/vintage-conversations.