Views
No views yet
SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: NewModel
(1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Dense({'in_features': 1024, 'out_features': 1024, 'bias': True, 'activation_function': 'torch.nn.modules.linear.Identity'})
)pip install -U sentence-transformers1from sentence_transformers import SentenceTransformer
2
3# Download from the 🤗 Hub
4model = SentenceTransformer("Gurveer05/stella-eedi-2024")
5# Run inference
6sentences = [
7 'Construct: Solve quadratic equations using the quadratic formula where the coefficient of x² is not 1.\n\nQuestion: Vera wants to solve this equation using the quadratic formula.\n(\n3 h^2-10 h+4=0\n)\n\nWhat should replace the circle? (? pm square root of (?-?) / bigcirc).\n\nOptions:\nA. 3\nB. 5\nC. 9\nD. 6\n\nCorrect Answer: 6\n\nIncorrect Answer: 3',
8 'Misremembers the quadratic formula',
9 'Does not know that vertically opposite angles are equal',
10]
11embeddings = model.encode(sentences)
12print(embeddings.shape)
13# [3, 1024]
14
15# Get the similarity scores for the embeddings
16similarities = model.similarity(embeddings, embeddings)
17print(similarities.shape)
18# [3, 3]qa_pair_text and MisconceptionName| qa_pair_text | MisconceptionName | |
|---|---|---|
| type | string | string |
| details |
|
|
| qa_pair_text | MisconceptionName |
|---|---|
Construct: Convert between cm³ and mm³.[object Object][object Object]Question: 1 cm^3 is the same as _______ mm^3.[object Object][object Object]Options:[object Object]A. 10[object Object]B. 100[object Object]C. 1000[object Object]D. 10000[object Object][object Object]Correct Answer: 1000[object Object][object Object]Incorrect Answer: 10 | Does not cube the conversion factor when converting cubed units |
Construct: Write algebraic expressions with correct algebraic convention.[object Object][object Object]Question: Which answer shows the following calculation using the correct algebraic convention?[object Object]([object Object]y x x+b x 3[object Object]).[object Object][object Object]Options:[object Object]A. y x+b 3[object Object]B. x y+3 b[object Object]C. y+3 b x[object Object]D. 3 b x y[object Object][object Object]Correct Answer: x y+3 b[object Object][object Object]Incorrect Answer: 3 b x y | Multiplies all terms together when simplifying an expression |
Construct: Write algebraic expressions with correct algebraic convention.[object Object][object Object]Question: Which of the following is the correct way of writing: p divided by q , then add 3 using algebraic convention?[object Object][object Object]Options:[object Object]A. p q+3[object Object]B. (p / q)+3[object Object]C. (p / q+3)[object Object]D. p-q+3[object Object][object Object]Correct Answer: (p / q)+3[object Object][object Object]Incorrect Answer: p-q+3 | Has used a subtraction sign to represent division |
MultipleNegativesRankingLoss with these parameters:
1{
2 "scale": 20.0,
3 "similarity_fct": "cos_sim"
4}qa_pair_text and MisconceptionName| qa_pair_text | MisconceptionName | |
|---|---|---|
| type | string | string |
| details |
|
|
| qa_pair_text | MisconceptionName |
|---|---|
Construct: Multiply two decimals together with the same number of decimal places.[object Object][object Object]Question: 0.4^2=.[object Object][object Object]Options:[object Object]A. 0.08[object Object]B. 0.8[object Object]C. 1.6[object Object]D. 0.16[object Object][object Object]Correct Answer: 0.16[object Object][object Object]Incorrect Answer: 0.8 | Mixes up squaring and multiplying by 2 or doubling |
Construct: Calculate the cube root of a number.[object Object][object Object]Question: 3rd root of (8)=.[object Object][object Object]Options:[object Object]A. 2 . dot{6}[object Object]B. 4[object Object]C. 64[object Object]D. 2[object Object][object Object]Correct Answer: 2[object Object][object Object]Incorrect Answer: 4 | Halves when asked to find the cube root |
Construct: Calculate missing lengths of shapes by geometrical inference, where the lengths given are in the same units.[object Object][object Object]Question: What is the area of the shaded section of this composite shape made from rectangles? A composite shape made from two rectangles that form an "L" shape. The base of the shape is horizontal and is 13cm long. The vertical height of the whole shape is 14cm. The horizontal width of the top part of the shape is 6cm. The vertical height of the top rectangle is 8cm. The right handed rectangle is shaded blue.[object Object][object Object]Options:[object Object]A. 48 cm^2[object Object]B. 104 cm^2[object Object]C. 42 cm^2[object Object]D. 56 cm^2[object Object][object Object]Correct Answer: 42 cm^2[object Object][object Object]Incorrect Answer: 48 cm^2 | Uses an incorrect side length when splitting a composite shape into parts |
MultipleNegativesRankingLoss with these parameters:
1{
2 "scale": 20.0,
3 "similarity_fct": "cos_sim"
4}eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16gradient_accumulation_steps: 32weight_decay: 0.01num_train_epochs: 20lr_scheduler_type: cosine_with_restartswarmup_ratio: 0.1fp16: Trueload_best_model_at_end: Truegradient_checkpointing: Truebatch_sampler: no_duplicatesoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 32eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.01adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 20max_steps: -1lr_scheduler_type: cosine_with_restartslr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Truefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Truegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseeval_use_gather_object: Falsebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportional| Epoch | Step | Training Loss | loss |
|---|---|---|---|
| 0.4183 | 2 | 0.9694 | - |
| 0.6275 | 3 | - | 0.8643 |
| 0.8366 | 4 | 0.8014 | - |
| 1.2549 | 6 | 0.5536 | 0.6725 |
| 1.6732 | 8 | 0.4849 | - |
| 1.8824 | 9 | - | 0.5312 |
| 2.0915 | 10 | 0.3331 | - |
| 2.5098 | 12 | 0.2363 | 0.4782 |
| 2.9281 | 14 | 0.1961 | - |
| 3.1373 | 15 | - | 0.4537 |
| 3.3464 | 16 | 0.1183 | - |
| 3.7647 | 18 | 0.0984 | 0.4637 |
| 4.1830 | 20 | 0.0617 | - |
| 4.3922 | 21 | - | 0.4658 |
1@inproceedings{reimers-2019-sentence-bert,
2 title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
3 author = "Reimers, Nils and Gurevych, Iryna",
4 booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
5 month = "11",
6 year = "2019",
7 publisher = "Association for Computational Linguistics",
8 url = "https://arxiv.org/abs/1908.10084",
9}1@misc{henderson2017efficient,
2 title={Efficient Natural Language Response Suggestion for Smart Reply},
3 author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
4 year={2017},
5 eprint={1705.00652},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}
8}