This dataset is part of the Anthropic's HH data used to train their RLHF Assistant
https://github.com/anthropics/hh-rlhf.
The data contains the first utterance from human to the dialog agent and the number of words in that utterance. The sampled version is a random sample of size 200.