Generated using QuicKB, a tool developed by Adam Lucek.
QuicKB optimizes document retrieval by creating fine-tuned knowledge bases through an end-to-end pipeline that handles document chunking, training data generation, and embedding model optimization.
Chunker: RecursiveTokenChunker
Parameters:
chunk_size: 400
chunk_overlap: 50
length_type: 'character'
separators: ['\n\n', '\n', '.', '?', '!', ' ', '']
keep_separator: True… See the full description on the dataset page:
https://huggingface.co/datasets/Mitchell6024/quickb.