Duplicated from philschmid/sharegpt-raw which is marked as duppliated from jeffwan/sharegpt_vicuna until step 3.
Then processed with the Better Uncensored (BUn) pipeline. A version with long conversations split is also provided.
Now we have the cleaned uncensored dataset in sharegpt_20230401_clean_bun.json (57058 conversations) and the same with split long conversations in sharegpt_20230401_clean_split_bun.json (103152 conversations).
This latest one should be a drop-in replacement for… See the full description on the dataset page:
https://huggingface.co/datasets/betteruncensored/sharegpt.