Filtered version of NilanE/ParallelFiction-Ja_En-100k in ShareGPT formatting and only uses a thousand examples.
I used Mistral Small 3.2 via OpenRouter to go through thousand examples to fix the translations. I haven't checked all 1000 examples but overlooked a bit and it seems fine but it's possible there are horrible issues with the dataset so I'd recommend still checking it manually if you want to be sure it's not trash.
Conversations do not exceed 16384 tokens which was based on Gemma 3… See the full description on the dataset page:
https://huggingface.co/datasets/mpasila/ParallelFiction-Ja_En-1k-16k-Gemma-3-ShareGPT-Filtered.