This is an experimental dataset to test a lower bound (resource wise) for generating paraphrases for scientific text in the chemistry domain. The goal is to see if we can generate paraphrases using a smaller model and less computational resources than typical frontier LLMs. As we also skip all quality filtering steps for the baseline, be extra careful if you use this dataset.
We use LM Studio with the "ibm/granite-4-h-tiny" (Q4_K_M GGUF) model to generate… See the full description on the dataset page:
https://huggingface.co/datasets/dwablimol/chemrxiv-paraphrase-demo.