Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("martintgc/finetuned_arctic_naive_ft-legal-ft-v0")
5# Run inference
6sentences = [
7 'Who did the captaincy of the New Zealand cricket team transition from, and how was this change perceived?',
8 "to be captain, but didn’t get it. Then, he kind of just took the captaincy from Ross Taylor in a bloodless coup. Although some people might say it wasn't particularly bloodless.And then New Zealand went to South Africa, and they had just the worst time. It was really, really bad. So all the players got together and asked themselves hard questions. “What are we doing? Are we going to become freelance players in franchise cricket? Are we going to take New Zealand cricket seriously?”Arising from that, they went all-in on a method that I suppose we would now call Bazball. But it did take a little while to come together. There were lots of different aspects that needed to fall in place. McCullum had to eventually give up the gloves, and most",
9 "were so good in this period is that they had two quality spinners, and they had a bunch of people who could play spin.By the time they got to white-ball sides also officially adopting Bazball, their ability to bowl spin and play spin had almost completely disappeared. They've tried to disrupt as much as possible, and they stole a Test in India doing it. But, we've also seen the limitations of that, and Pakistan embarrassed them recently. Now, they're back trying to play spin again, and we can see that they don't have it.The genius of Bazball and Brendon McCullum has always been to highlight your strengths and ignore your weaknesses. How do you do that when just playing spin or bowling spin is one of them? That's a hard thing to overcome",
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 1024]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]InformationRetrievalEvaluator| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.8571 |
| cosine_accuracy@3 | 0.9643 |
| cosine_accuracy@5 | 1.0 |
| cosine_accuracy@10 | 1.0 |
| cosine_precision@1 | 0.8571 |
| cosine_precision@3 | 0.3214 |
| cosine_precision@5 | 0.2 |
| cosine_precision@10 | 0.1 |
| cosine_recall@1 | 0.8571 |
| cosine_recall@3 | 0.9643 |
| cosine_recall@5 | 1.0 |
| cosine_recall@10 | 1.0 |
| cosine_ndcg@10 | 0.9354 |
| cosine_mrr@10 | 0.9137 |
| cosine_map@100 | 0.9137 |
sentence_0 and sentence_1| sentence_0 | sentence_1 | |
|---|---|---|
| type | string | string |
| details |
|
|
| sentence_0 | sentence_1 |
|---|---|
What is the main focus of Jarrod Kimber's "Bazball's mission to Moscow"? | Bazball's mission to Moscow - by Jarrod Kimber - Good Areas |
How does the concept of "Good Areas" relate to Bazball in the context of the article? | Bazball's mission to Moscow - by Jarrod Kimber - Good Areas |
What is the significance of Brendon McCullum and England’s Test side in relation to Bazball? | SubscribeSign inShare this postGood AreasBazball's mission to MoscowCopy linkFacebookEmailNotesMoreBazball's mission to MoscowIt created a revolution when Brendon McCullum and England’s Test side adopted it. But we’ve seen this before.Jarrod KimberFeb 13, 202513Share this postGood AreasBazball's mission to MoscowCopy linkFacebookEmailNotesMore1SharePre-order The Art of BattingSubscribeI've been thinking about Bazball a lot recently because Bazball’s back. Alright. We had Test Bazball, and now we have limited-overs Bazball. Which, I suppose, is Test Bazball on a slightly different frequency. But Bazball is much weirder when you think about it. I don’t know which version of Bazball we’re on, maybe it’s version three. Or maybe it’s version |
MatryoshkaLoss with these parameters:
1{
2 "loss": "MultipleNegativesRankingLoss",
3 "matryoshka_dims": [
4 768,
5 512,
6 256,
7 128,
8 64
9 ],
10 "matryoshka_weights": [
11 1,
12 1,
13 1,
14 1,
15 1
16 ],
17 "n_dims_per_step": -1
18}eval_strategy: stepsper_device_train_batch_size: 10per_device_eval_batch_size: 10num_train_epochs: 10multi_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 10per_device_eval_batch_size: 10per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 10max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin| Epoch | Step | cosine_ndcg@10 |
|---|---|---|
| 1.0 | 6 | 0.9232 |
| 2.0 | 12 | 0.9354 |
| 3.0 | 18 | 0.9486 |
| 4.0 | 24 | 0.9643 |
| 5.0 | 30 | 0.9643 |
| 6.0 | 36 | 0.9643 |
| 7.0 | 42 | 0.9643 |
| 8.0 | 48 | 0.9643 |
| 8.3333 | 50 | 0.9643 |
| 9.0 | 54 | 0.9643 |
| 10.0 | 60 | 0.9643 |
| 1.0 | 6 | 0.9643 |
| 2.0 | 12 | 0.9643 |
| 3.0 | 18 | 0.9379 |
| 4.0 | 24 | 0.9354 |
| 5.0 | 30 | 0.9354 |
| 6.0 | 36 | 0.9354 |
| 7.0 | 42 | 0.9339 |
| 8.0 | 48 | 0.9339 |
| 8.3333 | 50 | 0.9339 |
| 9.0 | 54 | 0.9339 |
| 10.0 | 60 | 0.9339 |
| 1.0 | 6 | 0.9486 |
| 2.0 | 12 | 0.9486 |
| 3.0 | 18 | 0.9486 |
| 4.0 | 24 | 0.9486 |
| 5.0 | 30 | 0.9486 |
| 6.0 | 36 | 0.9486 |
| 7.0 | 42 | 0.9354 |
| 8.0 | 48 | 0.9354 |
| 8.3333 | 50 | 0.9354 |
| 9.0 | 54 | 0.9354 |
| 10.0 | 60 | 0.9354 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{kusupati2024matryoshka,
2 title={Matryoshka Representation Learning},
3 author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
4 year={2024},
5 eprint={2205.13147},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG}
8}1@misc{henderson2017efficient,
2 title={Efficient Natural Language Response Suggestion for Smart Reply},
3 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
4 year={2017},
5 eprint={1705.00652},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}
8}