ArXiv_LLM is a fine-tuned version of all-mpnet-base-v2, trained on the ArXiv dataset. Due to computational limitations and restricted infrastructure, training was limited to just three epochs, preventing the model from achieving its full performance potential.
This model was developed as part of a Master's thesis in Applied Statistics and Data Science, titled "Optimizando la Búsqueda Semántica en ArXiv: Evaluación de Modelos LLMs y Herramientas de Indexación Vectorial"