Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 64, 'do_lower_case': False}) with Transformer model: NewModel
(1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("albertus-sussex/veriscrape-sbert-auto-wo-ref-deepseek-chat-0324")
5# Run inference
6sentences = [
7 '$27,174',
8 'The data provided by Autodata is provided AS IS without warranty or guarantee of any kind, and Autodata disclaims all warranties or conditions of any kind, expressed or implied, with respect to such data, including the implied warranties of merchantable quality and fitness for a particular purpose.',
9 '$39,890',
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 768]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]TripletEvaluator| Metric | Value |
|---|---|
| cosine_accuracy | 0.9837 |
veriscrape.training.SilhouetteEvaluator| Metric | Value |
|---|---|
| silhouette_cosine | 0.405 |
| silhouette_euclidean | 0.321 |
TripletEvaluator| Metric | Value |
|---|---|
| cosine_accuracy | 0.9793 |
veriscrape.training.SilhouetteEvaluator| Metric | Value |
|---|---|
| silhouette_cosine | 0.4049 |
| silhouette_euclidean | 0.3216 |
anchor, positive, negative, pos_attr_name, neg_attr_name, and website_id| anchor | positive | negative | pos_attr_name | neg_attr_name | website_id | |
|---|---|---|---|---|---|---|
| type | string | string | string | string | string | int |
| details |
|
|
|
|
|
|
| anchor | positive | negative | pos_attr_name | neg_attr_name | website_id |
|---|---|---|---|---|---|
$34,270 | $22,240 | Lexus GX 460 Base 4dr AWD | price | model | 0 |
FWD or AWD | - | GT-R | engine | model | 5 |
- | $15,195 | City: [object Object] [object Object] 11 [object Object] [object Object] [object Object] [object Object] [object Object] Highway: [object Object] [object Object] 17[object Object] – [object Object] [object Object] 18 | engine | fuel_economy | 5 |
TripletLoss with these parameters:
1{
2 "distance_metric": "TripletDistanceMetric.EUCLIDEAN",
3 "triplet_margin": 5
4}anchor, positive, negative, pos_attr_name, neg_attr_name, and website_id| anchor | positive | negative | pos_attr_name | neg_attr_name | website_id | |
|---|---|---|---|---|---|---|
| type | string | string | string | string | string | int |
| details |
|
|
|
|
|
|
| anchor | positive | negative | pos_attr_name | neg_attr_name | website_id |
|---|---|---|---|---|---|
$23,215 | $95,465 | $245,000 | engine | price | 5 |
Visit our partners: | | $32,000 – $38,000 | engine | price | |
$44,605 | $22,530 | 22 mpg city / 33 mpg hwy | price | fuel_economy | 9 |
TripletLoss with these parameters:
1{
2 "distance_metric": "TripletDistanceMetric.EUCLIDEAN",
3 "triplet_margin": 5
4}eval_strategy: epochper_device_train_batch_size: 128per_device_eval_batch_size: 128num_train_epochs: 5warmup_ratio: 0.1overwrite_output_dir: Falsedo_predict: Falseeval_strategy: epochprediction_loss_only: Trueper_device_train_batch_size: 128per_device_eval_batch_size: 128per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 5max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseeval_use_gather_object: Falseprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | Validation Loss | cosine_accuracy | silhouette_cosine |
|---|---|---|---|---|---|
| -1 | -1 | - | - | 0.5342 | 0.1323 |
| 1.0 | 276 | 0.5187 | 0.2374 | 0.9829 | 0.3699 |
| 2.0 | 552 | 0.0959 | 0.1561 | 0.9880 | 0.3959 |
| 3.0 | 828 | 0.0714 | 0.1738 | 0.9878 | 0.4028 |
| 4.0 | 1104 | 0.0594 | 0.1711 | 0.9875 | 0.4159 |
| 5.0 | 1380 | 0.05 | 0.2089 | 0.9837 | 0.4050 |
| -1 | -1 | - | - | 0.9793 | 0.4049 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{hermans2017defense,
2 title={In Defense of the Triplet Loss for Person Re-Identification},
3 author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
4 year={2017},
5 eprint={1703.07737},
6 archivePrefix={arXiv},
7 primaryClass={cs.CV}
8}