Lemone-embedded, pre-built embeddings dataset for French taxation.
This database presents the embeddings generated by the Lemone-embed-pro model and aims at a large-scale distribution of the model even for the GPU-poor.
This sentence transformers model, specifically designed for French taxation, has been fine-tuned on a dataset comprising 43 million tokens, integrating a blend of semi-synthetic and fully synthetic data generated by GPT-4 Turbo and Llama 3.1 70B, which have… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/lemone-docs-embedded.