Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 8192, 'do_lower_case': False}) with Transformer model: ModernBertModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("kokojake/modernbert-embed-base-fitness-health-matryoshka")
5# Run inference
6sentences = [
7 'facilitate publication;\n•\u2009\x07Mobilise academic expertise for \ndeveloping training programmes and \nmobilising trainers.\n\t Weigh in on the debate around issues \nrelated to rehabilitation promotion and funding, promote best practices to \ninfluence policies that favour access \nto rehabilitation services and thereby \nmove toward advocacy actions.\n48\nUsers,\nDisabled people’s\norganisations\nService\nproviders\nDecision-makers User \ngroups\nLocal\nauthorities\nMinistry of \nHealth, Ministry \nof Social Action,\netc.\nUnited Nations \n(WHO, etc.)\nHospitals, \nReference\nrehabilitation centre\nProfessional \nassociations\nService provider groups\nTraining institutes\nCommunity- \nbased Services\nFederation\nand national\n associations\nHospital, \nHealth \ncare centres Network: actors that can be mobilised for physical \nand functional rehabilitation\nInternational\nNational\nLocal\nInstitutional donors\nFacilitation organisations* * \x07Organisations (IOs, NGOs, etc.), agencies, universities and research centres that facilitate the existence of physical \nand functional rehabilitation via national or international projects.\nInternational \n consortia (IDDC, etc.)\n International\n networks \n (CBR, WCPT, \n WFOT, ISPO,\nFATO, etc.)\nLevels of intervention © Handicap International, 2013\n \n \n49\n\xa0Intervention.\n\xa0modalities\u200a.\nThe Unit has technical resources specifically \npositioned to be able to reach the maximum',
8 'training programmes for rehabilitation professionals',
9 'risks of yo-yo dieting and heart disease',
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 768]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]dim_768InformationRetrievalEvaluator with these parameters:
1{
2 "truncate_dim": 768
3}| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.4789 |
| cosine_accuracy@3 | 0.4789 |
| cosine_accuracy@5 | 0.4789 |
| cosine_accuracy@10 | 0.5219 |
| cosine_precision@1 | 0.4789 |
| cosine_precision@3 | 0.4789 |
| cosine_precision@5 | 0.4789 |
| cosine_precision@10 | 0.4395 |
| cosine_recall@1 | 0.0601 |
| cosine_recall@3 | 0.1802 |
| cosine_recall@5 | 0.3003 |
| cosine_recall@10 | 0.5134 |
| cosine_ndcg@10 | 0.499 |
| cosine_mrr@10 | 0.4861 |
| cosine_map@100 | 0.5681 |
dim_512InformationRetrievalEvaluator with these parameters:
1{
2 "truncate_dim": 512
3}| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.4742 |
| cosine_accuracy@3 | 0.4742 |
| cosine_accuracy@5 | 0.4742 |
| cosine_accuracy@10 | 0.5141 |
| cosine_precision@1 | 0.4742 |
| cosine_precision@3 | 0.4742 |
| cosine_precision@5 | 0.4742 |
| cosine_precision@10 | 0.4362 |
| cosine_recall@1 | 0.059 |
| cosine_recall@3 | 0.1769 |
| cosine_recall@5 | 0.2948 |
| cosine_recall@10 | 0.5078 |
| cosine_ndcg@10 | 0.4934 |
| cosine_mrr@10 | 0.4808 |
| cosine_map@100 | 0.5632 |
dim_256InformationRetrievalEvaluator with these parameters:
1{
2 "truncate_dim": 256
3}| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.4555 |
| cosine_accuracy@3 | 0.4555 |
| cosine_accuracy@5 | 0.4555 |
| cosine_accuracy@10 | 0.4969 |
| cosine_precision@1 | 0.4555 |
| cosine_precision@3 | 0.4555 |
| cosine_precision@5 | 0.4555 |
| cosine_precision@10 | 0.4188 |
| cosine_recall@1 | 0.057 |
| cosine_recall@3 | 0.1711 |
| cosine_recall@5 | 0.2851 |
| cosine_recall@10 | 0.4882 |
| cosine_ndcg@10 | 0.4746 |
| cosine_mrr@10 | 0.4624 |
| cosine_map@100 | 0.5446 |
dim_128InformationRetrievalEvaluator with these parameters:
1{
2 "truncate_dim": 128
3}| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.4352 |
| cosine_accuracy@3 | 0.4352 |
| cosine_accuracy@5 | 0.4352 |
| cosine_accuracy@10 | 0.4727 |
| cosine_precision@1 | 0.4352 |
| cosine_precision@3 | 0.4352 |
| cosine_precision@5 | 0.4352 |
| cosine_precision@10 | 0.3988 |
| cosine_recall@1 | 0.0545 |
| cosine_recall@3 | 0.1636 |
| cosine_recall@5 | 0.2726 |
| cosine_recall@10 | 0.4639 |
| cosine_ndcg@10 | 0.4522 |
| cosine_mrr@10 | 0.4414 |
| cosine_map@100 | 0.5208 |
dim_64InformationRetrievalEvaluator with these parameters:
1{
2 "truncate_dim": 64
3}| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.3945 |
| cosine_accuracy@3 | 0.3945 |
| cosine_accuracy@5 | 0.3945 |
| cosine_accuracy@10 | 0.4297 |
| cosine_precision@1 | 0.3945 |
| cosine_precision@3 | 0.3945 |
| cosine_precision@5 | 0.3945 |
| cosine_precision@10 | 0.3598 |
| cosine_recall@1 | 0.0499 |
| cosine_recall@3 | 0.1497 |
| cosine_recall@5 | 0.2495 |
| cosine_recall@10 | 0.4224 |
| cosine_ndcg@10 | 0.4109 |
| cosine_mrr@10 | 0.4004 |
| cosine_map@100 | 0.4763 |
positive and anchor| positive | anchor | |
|---|---|---|
| type | string | string |
| details |
|
|
| positive | anchor |
|---|---|
values and preferences among older people [object Object]in relation to exercise, noting that older [object Object]people valued the outcomes of exercise [object Object]for maintaining health. They judged that [object Object]the evidence for older people was likely to be relevant to all adults and agreed [object Object]there was likely to be some uncertainty or [object Object]variability with respect to people’s values and [object Object]preferences for exercise and its outcomes. [object Object]Some GDG members suggested that given reasonably consistent benefit and very [object Object]little harms, there would be no important [object Object]uncertainty or variability regarding people’s [object Object]values on the outcomes of exercise. In the [object Object]absence of direct qualitative evidence, the GDG judged from their own experience [object Object]that resource requirements for structured [object Object]exercise programmes would vary by country [object Object]and setting, but in some settings might [object Object]be associated with moderate costs (for [object Object]structured exercise programmes, compared with self-managed physical activity). The GDG [object Object]noted that costs could also vary according to [object Object]the modality of ... | exercise preferences and outcomes variability among adults |
ICRC, ICRC Hospital Design and Rehabilitation Guidelines, Vol. 1: Models Of Care, ICRC, Geneva, 2022: [object Object]. icrc.org/icrc-hospital-design-and-rehabilitation-guidelines-volume-1-models-of-care-print-en.html | ICRC rehabilitation guidelines 2022 |
fitness training is guided by a health worker or (if feasible) performed self-directed by [object Object]the patient following education and advice.[object Object]Metacognitive [object Object]training[object Object]Metacognitive training aims to improve social functioning through reducing cognitive biases/psychotic symptoms (e.g. delusion, impaired self-awareness or insight). [object Object]Metacognitive training is usually provided as a structured group intervention during which participants perform exercises to reflect their own thinking and receive training in [object Object]strategies to cope with cognitive biases during daily routines. Metacognitive training is [object Object]guided by a health worker.[object Object]Mindfulness-[object Object]based approaches Mindfulness-based interventions aim to achieve a state of mindfulness in which a [object Object]person becomes more aware of their physical, mental, and emotional condition in the [object Object]present moment, without becoming judgemental. Mindfulness-based interventions (e.g. mindfulness-based cognitive therapy, acceptance and commitment therapy) [object Object]help people to pay attentio... | structured group interventions for metacognitive training |
MatryoshkaLoss with these parameters:
1{
2 "loss": "MultipleNegativesRankingLoss",
3 "matryoshka_dims": [
4 768,
5 512,
6 256,
7 128,
8 64
9 ],
10 "matryoshka_weights": [
11 1,
12 1,
13 1,
14 1,
15 1
16 ],
17 "n_dims_per_step": -1
18}eval_strategy: epochper_device_train_batch_size: 32per_device_eval_batch_size: 16gradient_accumulation_steps: 16learning_rate: 2e-05num_train_epochs: 4lr_scheduler_type: cosinewarmup_ratio: 0.1bf16: Truetf32: Trueload_best_model_at_end: Trueoptim: adamw_torch_fusedbatch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: epochprediction_loss_only: Trueper_device_train_batch_size: 32per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 16eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 2e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 4max_steps: -1lr_scheduler_type: cosinelr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Truefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Truelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}tp_size: 0fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torch_fusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | dim_768_cosine_ndcg@10 | dim_512_cosine_ndcg@10 | dim_256_cosine_ndcg@10 | dim_128_cosine_ndcg@10 | dim_64_cosine_ndcg@10 |
|---|---|---|---|---|---|---|---|
| 0.4444 | 10 | 64.4729 | - | - | - | - | - |
| 0.8889 | 20 | 32.1029 | - | - | - | - | - |
| 1.0 | 23 | - | 0.4734 | 0.4741 | 0.4590 | 0.4271 | 0.3722 |
| 1.3111 | 30 | 23.9454 | - | - | - | - | - |
| 1.7556 | 40 | 19.7319 | - | - | - | - | - |
| 2.0 | 46 | - | 0.4934 | 0.4926 | 0.4723 | 0.4471 | 0.4021 |
| 2.1778 | 50 | 17.6381 | - | - | - | - | - |
| 2.6222 | 60 | 16.9329 | - | - | - | - | - |
| 3.0 | 69 | - | 0.498 | 0.4954 | 0.4746 | 0.4528 | 0.4089 |
| 3.0444 | 70 | 15.4096 | - | - | - | - | - |
| 3.4889 | 80 | 15.4012 | - | - | - | - | - |
| 3.8444 | 88 | - | 0.4990 | 0.4934 | 0.4746 | 0.4522 | 0.4109 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{kusupati2024matryoshka,
2 title={Matryoshka Representation Learning},
3 author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
4 year={2024},
5 eprint={2205.13147},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG}
8}1@misc{henderson2017efficient,
2 title={Efficient Natural Language Response Suggestion for Smart Reply},
3 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
4 year={2017},
5 eprint={1705.00652},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}
8}