This dataset is a filtered subset of neo4j/text2cypher-2024v1.
Filtered to examples where len(question) + len(schema) + len(cypher) < 1500 characters
train split: 1000 examples (shuffled with seed=42, then sliced)
val split: 75 examples (same shuffle/slice procedure)
All other fields preserved (question, schema, cypher, data_source, ...)
Apache License 2.0 (inherited… See the full description on the dataset page:
https://huggingface.co/datasets/RomanTeucher/text2cypher-curated.