Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("chelleboyer/llm-mm-good-309e6f79-505b-4c23-8452-37cc854e67df")
5# Run inference
6sentences = [
7 'What metrics does the LLMS (2025) framework introduce to investigate position bias in pairwise comparisons?',
8 'Recent studies have further examined position bias in the LLMs-as-judges context.\nFor instance, a framework\xa0(LLMS, 2025) is proposed to investigate position bias in pairwise comparisons, introducing metrics such as repetition stability, position consistency, and preference fairness to better understand how positions affect LLM judgments.\nAnother study\xa0(Zheng et\xa0al., 2023a) explores the limitations of LLMs-as-judges, including position biases, and verifies agreement between LLM judgments and human preferences across multiple benchmarks.\nThese findings underscore the need for robust debiasing strategies to enhance the fairness and reliableness of LLMs-as-judges.',
9 'Overconfidence bias\xa0(Khan et\xa0al., 2024; Jung et\xa0al., 2024) in the context of LLMs-as-judges refers to the tendency of models to exhibit an inflated level of confidence in their judgments, often resulting in overly assertive evaluations that may not accurately reflect the true reliability of the answer. This bias is particularly concerning in evaluative contexts, as it can lead LLMs-as-judges to overstate the correctness of certain outputs, compromising the objectivity and dependability of assessments.',
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 1024]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]InformationRetrievalEvaluator| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.93 |
| cosine_accuracy@3 | 0.99 |
| cosine_accuracy@5 | 1.0 |
| cosine_accuracy@10 | 1.0 |
| cosine_precision@1 | 0.93 |
| cosine_precision@3 | 0.33 |
| cosine_precision@5 | 0.2 |
| cosine_precision@10 | 0.1 |
| cosine_recall@1 | 0.93 |
| cosine_recall@3 | 0.99 |
| cosine_recall@5 | 1.0 |
| cosine_recall@10 | 1.0 |
| cosine_ndcg@10 | 0.9704 |
| cosine_mrr@10 | 0.9603 |
| cosine_map@100 | 0.9603 |
sentence_0 and sentence_1| sentence_0 | sentence_1 | |
|---|---|---|
| type | string | string |
| details |
|
|
| sentence_0 | sentence_1 |
|---|---|
What are the main components of the evaluation function ( E ) as described in the preliminaries section? | LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods[object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object]1 Introduction[object Object][object Object]2 PRELIMINARIES[object Object][object Object]2.1 Evaluation Function E𝐸Eitalic_E[object Object][object Object]2.2 Evaluation Input[object Object][object Object]2.2.1 Evaluation Type 𝒯𝒯\mathcal{T}caligraphic_T[object Object]2.2.2 Evaluation Criteria 𝒞𝒞\mathcal{C}caligraphic_C.[object Object]2.2.3 Evaluation References ℛℛ\mathcal{R}caligraphic_R.[object Object][object Object][object Object]2.3 Evaluation Output[object Object][object Object][object Object][object Object]3 Functionality[object Object][object Object][object Object]3.1 Performance Evaluation[object Object][object Object]3.1.1 Responses Evaluation[object Object]3.1.2 Model Evaluation[object Object][object Object][object Object][object Object]3.2 Model Enhancement[object Object][object Object]3.2.1 Reward Modeling During Training[object Object]3.2.2 Acting as Verifier During Inference[object Object]3.2.3 Feedback for Refinement[object Object][object Object][object Object][object Object]3.3 Data Construction[object Object][object Object]3.3.1 Data Annotation[object Object]3.3.2 Data Synthesize[object Object][object Object][object Object][object Object][object Object][object Object]4 Methodology |
How do LLMs contribute to model enhancement according to the functionalities outlined in the survey? | LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods[object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object][object Object]1 Introduction[object Object][object Object]2 PRELIMINARIES[object Object][object Object]2.1 Evaluation Function E𝐸Eitalic_E[object Object][object Object]2.2 Evaluation Input[object Object][object Object]2.2.1 Evaluation Type 𝒯𝒯\mathcal{T}caligraphic_T[object Object]2.2.2 Evaluation Criteria 𝒞𝒞\mathcal{C}caligraphic_C.[object Object]2.2.3 Evaluation References ℛℛ\mathcal{R}caligraphic_R.[object Object][object Object][object Object]2.3 Evaluation Output[object Object][object Object][object Object][object Object]3 Functionality[object Object][object Object][object Object]3.1 Performance Evaluation[object Object][object Object]3.1.1 Responses Evaluation[object Object]3.1.2 Model Evaluation[object Object][object Object][object Object][object Object]3.2 Model Enhancement[object Object][object Object]3.2.1 Reward Modeling During Training[object Object]3.2.2 Acting as Verifier During Inference[object Object]3.2.3 Feedback for Refinement[object Object][object Object][object Object][object Object]3.3 Data Construction[object Object][object Object]3.3.1 Data Annotation[object Object]3.3.2 Data Synthesize[object Object][object Object][object Object][object Object][object Object][object Object]4 Methodology |
What are the different approaches discussed under the Single-LLM System methodology? | 4 Methodology[object Object][object Object][object Object]4.1 Single-LLM System[object Object][object Object]4.1.1 Prompt-based[object Object]4.1.2 Tuning-based[object Object]4.1.3 Post-processing[object Object][object Object][object Object][object Object]4.2 Multi-LLM System[object Object][object Object]4.2.1 Communication[object Object]4.2.2 Aggregation[object Object][object Object][object Object]4.3 Human-AI Collaboration System[object Object][object Object][object Object][object Object]5 Application[object Object][object Object]5.1 General[object Object]5.2 Multimodal[object Object]5.3 Medical[object Object]5.4 Legal[object Object]5.5 Financial[object Object]5.6 Education[object Object]5.7 Information Retrieval[object Object][object Object]5.8 Others[object Object][object Object]5.8.1 Soft Engineering[object Object]5.8.2 Biology[object Object]5.8.3 Social Science[object Object][object Object][object Object][object Object][object Object][object Object]6 Meta-evaluation[object Object][object Object][object Object]6.1 Benchmarks[object Object][object Object]6.1.1 Code Generation[object Object]6.1.2 Machine Translation[object Object]6.1.3 Text Summarization[object Object]6.1.4 Dialogue Generation[object Object]6.1.5 Automatic Story Generation[object Object]6.1.6 Values Alignment[object Object]6.1.7 Recommendation[object Object]6.1.8 Search[object Object]6.1.9 Comprehensive Data[object Object][object Object][object Object][object Object]6.2 Metric |
MatryoshkaLoss with these parameters:
1{
2 "loss": "MultipleNegativesRankingLoss",
3 "matryoshka_dims": [
4 768,
5 512,
6 256,
7 128,
8 64
9 ],
10 "matryoshka_weights": [
11 1,
12 1,
13 1,
14 1,
15 1
16 ],
17 "n_dims_per_step": -1
18}eval_strategy: stepsper_device_train_batch_size: 50per_device_eval_batch_size: 50num_train_epochs: 10multi_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 50per_device_eval_batch_size: 50per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 10max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}tp_size: 0fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseaverage_tokens_across_devices: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin| Epoch | Step | cosine_ndcg@10 |
|---|---|---|
| 1.0 | 27 | 0.9697 |
| 1.8519 | 50 | 0.9788 |
| 2.0 | 54 | 0.9775 |
| 3.0 | 81 | 0.9741 |
| 3.7037 | 100 | 0.9791 |
| 4.0 | 108 | 0.9741 |
| 5.0 | 135 | 0.9782 |
| 5.5556 | 150 | 0.9782 |
| 6.0 | 162 | 0.9782 |
| 7.0 | 189 | 0.9782 |
| 7.4074 | 200 | 0.9741 |
| 8.0 | 216 | 0.9741 |
| 9.0 | 243 | 0.9704 |
| 9.2593 | 250 | 0.9704 |
| 10.0 | 270 | 0.9704 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{kusupati2024matryoshka,
2 title={Matryoshka Representation Learning},
3 author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
4 year={2024},
5 eprint={2205.13147},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG}
8}1@misc{henderson2017efficient,
2 title={Efficient Natural Language Response Suggestion for Smart Reply},
3 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
4 year={2017},
5 eprint={1705.00652},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}
8}