One 768-d float16 vector per utterance of allenai/soda, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
soda_e_t.f16.npy — numpy array (n_turns, 768), float16;… See the full description on the dataset page:
https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-soda.