Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("mazej/all-MiniLM-L6-v2_MultipleNegativesRankingLoss_b64_accs4_171531_fine-tuned")
5# Run inference
6sentences = [
7 'Are files in a hidden directory safe?',
8 "Are files in a hidden directory safe? Lets pretend that I have a website and want to keep personal notes on it. If I make a directory like mysite.com/84eW93e7 or mysite.com/34aWe897.txt where the letters and numbers are random so randomly guessing the file or directory name are impractical. How secure is this directory? Could someone still discover the existence of the file or directory assuming directory listing is off? If so, how would they find out? Of course this isnt something I would do, I'm just curious.",
9 'I have used the ogr2ogr to import various formats to Postgres and have used GDAL utilities. Is there a complete Java port for these tools? Or are they simply wrappers calling JNI to C++ code? I have seen some Java stuff by googling but not sure how much are they current and kept upto date. Thanks, Ramesh',
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 384]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]ir_evalInformationRetrievalEvaluator| Metric | Value |
|---|---|
| cosine_accuracy@10 | 0.7569 |
| cosine_precision@10 | 0.1019 |
| cosine_recall@10 | 0.705 |
| cosine_ndcg@10 | 0.595 |
| cosine_mrr@10 | 0.5857 |
| cosine_map@100 | 0.5534 |
| dot_accuracy@10 | 0.7569 |
| dot_precision@10 | 0.1019 |
| dot_recall@10 | 0.705 |
| dot_ndcg@10 | 0.595 |
| dot_mrr@10 | 0.5857 |
| dot_map@100 | 0.5534 |
query and document| query | document | |
|---|---|---|
| type | string | string |
| details |
|
|
| query | document |
|---|---|
Differnce between abilities of System apps, Apps in phone memory and Apps in the SD card | As I've played around with different ROMs and used Titanium Backup, I see that apps are categorized into system and user apps. Several ROM developers state that Titanium Backup should only be used for user apps and not for system apps, and when I proceed to uninstall system apps, TiBu warns me that the ROM may not work correctly. Other than the fact that user apps are downloadable via the Market or some other means, what are the key differences between the two? What is it about the system apps that make them more integral to the OS itself? |
Usage of Who and Whom | What's the rule for using who or whom? I was writing a LinkedIn recommendation one day, and ended up pondering for a while which of these forms to use: > … is a great developer whom I always found easy to work with. > > … is a great developer who I always found easy to work with. Both are basically correct in contemporary English, right? But is one or the other preferable, and if so, why? (In this case I went with the latter, as it seemed more common (Google) and I wanted to avoid sounding unnecessarily “archaic”, although I’m not sure whether that would have been the case. Also note that I try to write in a “friendly professional” style instead of overly formal one. :-) |
How do I override a \configure statement in html4.4t - tex4ht | Im trying to prepare html for conversion to an epub and have encountered a couple of problems that I would appreciate help with. Hope mixing a couple of related questions in a single post is OK. The easiest is the quotation environment doesnt indent both margins (and not enough on the left). Ive written a .cfg file and think I have that one solved. Heres what Im using: \Preamble{html} \begin{document} \Css { .quotation { margin-bottom:0.25em; margin-top:0.25em; margin-left:4em; margin-right:4em } } \EndPreamble The next is figures and tables. tex4ht inserts an hr tag before and after the table. This is not part of the css, its a separate tag generated by tex4ht. I havent found where one customizes this. I can delete these in emacs from the html, but that doesnt seem the right way to do this. Heres a fairly minimal example showing the rules (and the quotation indent as well). \documentclass{book} \begin{document} \chapter{Simple Example} Lorem ipsum dolor sit amet, consectetur adipiscing elit. Integer nec odio. Praesent libero. Sed cursus ante dapibus diam. Sed nisi. Nulla quis sem at nibh elementum imperdiet. Duis sagittis ipsum. Praesent mauris. Fusce nec tellus sed augue semper porta. Mauris massa. Vestibulum lacinia arcu eget nulla. \begin{quotation} Default book style icreases the left and right margins for quotations by a noticeable amount. The css generated by tex4ht only increases the left margin by one em. The attached .cfg increases this to 4 em. \end{quotation} Sed lectus. Integer euismod lacus luctus magna. Quisque cursus, metus vitae pharetra auctor, sem massa mattis sem, at interdum magna augue eget diam. Vestibulum ante ipsum primis in faucibus orci luctus et ultrices posuere cubilia Curae; Morbi lacinia molestie dui. amet mauris. \begin{table} \begin{center} \hbox{ \begin{tabular}{ |
MultipleNegativesRankingLoss with these parameters:
1{
2 "scale": 20.0,
3 "similarity_fct": "cos_sim"
4}per_device_train_batch_size: 64per_device_eval_batch_size: 64gradient_accumulation_steps: 4warmup_ratio: 0.1batch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: noprediction_loss_only: Trueper_device_train_batch_size: 64per_device_eval_batch_size: 64per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 4eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 3max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falsebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | ir_eval_cosine_map@100 |
|---|---|---|
| 2.8276 | 63 | 0.5534 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{henderson2017efficient,
2 title={Efficient Natural Language Response Suggestion for Smart Reply},
3 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
4 year={2017},
5 eprint={1705.00652},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}
8}