This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search.
We follow a teacher-student transfer learning approach to train an bertin-roberta-base-spanish model using parallel EN-ES sentence pairs.
1from sentence_transformers import SentenceTransformer
2sentences =["Este es un ejemplo","Cada oración es transformada"]34model = SentenceTransformer('hackathon-pln-es/paraphrase-spanish-distilroberta')5embeddings = model.encode(sentences)6print(embeddings)
Usage (HuggingFace Transformers)
Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.
python
1from transformers import AutoTokenizer, AutoModel
2import torch
3import torch.nn.functional as F
45#Mean Pooling - Take attention mask into account for correct averaging6defmean_pooling(model_output, attention_mask):7 token_embeddings = model_output[0]#First element of model_output contains all token embeddings8 input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()9return torch.sum(token_embeddings * input_mask_expanded,1)/ torch.clamp(input_mask_expanded.sum(1),min=1e-9)101112# Sentences we want sentence embeddings for13sentences =['Este es un ejemplo", "Cada oración es transformada']1415# Load model from HuggingFace Hub16tokenizer = AutoTokenizer.from_pretrained('hackathon-pln-es/paraphrase-spanish-distilroberta')17model = AutoModel.from_pretrained('hackathon-pln-es/paraphrase-spanish-distilroberta')1819# Tokenize sentences20encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')2122# Compute token embeddings23with torch.no_grad():24 model_output = model(**encoded_input)2526# Perform pooling27sentence_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])2829# Normalize embeddings30sentence_embeddings = F.normalize(sentence_embeddings, p=2, dim=1)3132print("Sentence embeddings:")33print(sentence_embeddings)
Similarity Evaluation on STS-2017.es-en.txt and STS-2017.es-es.txt (translated manually for evaluation purposes)
We measure the semantic textual similarity (STS) between sentence pairs in different languages:
ES-ES
cosine_pearson
cosine_spearman
manhattan_pearson
manhattan_spearman
euclidean_pearson
euclidean_spearman
dot_pearson
dot_spearman
0.8495
0.8579
0.8675
0.8474
0.8676
0.8478
0.8277
0.8258
ES-EN
cosine_pearson
cosine_spearman
manhattan_pearson
manhattan_spearman
euclidean_pearson
euclidean_spearman
dot_pearson
dot_spearman
0.8344
0.8448
0.8279
0.8168
0.8282
0.8159
0.8083
0.8145
Intended uses
Our model is intented to be used as a sentence and short paragraph encoder. Given an input text, it ouptuts a vector which captures
the semantic information. The sentence vector may be used for information retrieval, clustering or sentence similarity tasks.
We use the concatenation from multiple datasets with sentence pairs (EN-ES).
We could check out the dataset that was used during training: parallel-sentences