A mechanism tutor that diagnoses why an answer is right or wrong, not merely which product appears at the end.
LoRA-adapted from Llama 3.3 70B Instruct with a 10,000-row, ground-truth-separated training set.
BondShift evaluation results
TL;DR
BondShift targets a common failure in organic-chemistry assistance: producing a plausible product while giving an
invalid electron-flow story. It connects reagents and conditions to electron movement, charge, intermediates,
stereochemical constraints, pathway choice, and the final outcome.
The target behavior is diagnostic. BondShift should identify the first invalid step, explain the controlling chemical
principle, repair the mechanism, and state what additional structural or condition information is needed when the answer
is genuinely underdetermined.
The submitted model is a LoRA adapter, not a standalone 70B checkpoint. It must be used with the exact base-model
family recorded in the AutoScientist configuration.
Organic chemistry is not solved by reaction-name recall alone. A student can memorize "strong base means E2" and still
draw an impossible arrow, remove the wrong beta hydrogen, miss an anti-periplanar requirement, create an unstable
carbocation, or treat resonance contributors as rapidly interconverting molecules.
These are high-value tutoring failures because the final product can look correct even when the reasoning that produced
it is not. Conventional answer-only data rewards the destination and may never teach the model to locate the broken
step.
BondShift therefore trains the reasoning layer between problem and conclusion:
Failure mode
Desired BondShift behavior
Correct product, invalid mechanism
Find and repair the first chemically invalid step
Mechanism chosen from one keyword
Weigh substrate, nucleophile/base, solvent, geometry, and conditions together
Strong reagent treated as overriding structure
Explain the geometric or orbital constraint that still applies
Missing structure or conditions
Give a bounded answer or ask for the decisive missing fact
Resonance or charge misconception
Track electron and charge conservation explicitly
Unsupported selectivity claim
Avoid false precision and state the decision-changing conditions
The design response
BondShift was built around independently constructed chemistry answer keys instead of allowing one language model to
invent the question, answer, and validation truth together. The response target is still natural tutoring prose: answer
first, explain the decisive mechanism, correct the misconception, and add caveats only when they matter.
The production design also deliberately restricted V1 to chemistry families whose ground truth could be represented
reliably. Breadth was postponed when it would have required fragile or overconfident answer keys.
Supported capabilities
Capability
Examples
Mechanism debugging
Find an invalid arrow, impossible intermediate, or wrong leaving-group step
Outcome reasoning
Predict a qualitative product or pathway from a supported substrate and conditions
Key-step explanation
Explain proton transfer, ionization, backside attack, or beta elimination
Pathway choice
Compare supported SN1, SN2, E1, and E2 pathways
Stereochemical checks
Explain inversion and cyclohexane trans-diaxial requirements when the structure is sufficiently specified
Misconception repair
Correct common errors about charge, resonance, nucleophilicity, basicity, and carbocations
Science bridge explanations
Connect arrow pushing to charge flow, orbital alignment, and observable chemical behavior
The validated V1 training scope contains six families: acid-base, resonance/formal charge, SN1,
SN2, E1, and E2.
Example
This is an actual prompt-response pair from the frozen training snapshot.
User
I'm looking at this E2 with bromocyclohexane and NaOEt/EtOH heat. My notes say any beta-H can go because ethoxide is
strong, but the problem hints the leaving group has to be axial and a chair flip might be needed. Does that mean only
certain hydrogens actually work, or does the strong base override the geometry thing?
BondShift
In an E2 reaction, the base strength doesn't override the geometric requirement. The reaction is concerted and
demands that the breaking C-H and C-Br bonds be anti-periplanar so their sigma orbitals can align with the forming pi
bond. In a cyclohexane chair, this alignment is only achieved when the leaving group (Br) and a beta-hydrogen are both
axial, i.e. trans-diaxial.
If bromocyclohexane starts with Br equatorial, a chair flip is needed to place Br axial. Only then can an axial
beta-hydrogen on an adjacent carbon participate. Hydrogens that are equatorial, or axial but not anti to the leaving
group, cannot react because the orbital overlap is wrong. So, despite ethoxide being a strong base, only the
trans-diaxial hydrogens are geometrically competent, and the product is cyclohexene.
Evaluation
AutoScientist win rates
Evaluation slice
Base model
Adapted model
AutoScientist evaluation on this dataset
3
98
Science category evaluation
14
86
These are the whole-number win-rate labels displayed by the Adaption AutoScientist evaluation interface for training
experiment c7f0a1b0-8286-4387-8f08-f2e0a4b74998. The interface did not expose sample counts, confidence intervals,
or a public item-level evaluation set. The values should therefore be read as platform-reported preference results,
not as universal chemistry accuracy estimates. The own-dataset labels may also reflect display rounding.
Dataset adaptation signals
The dataset used for this run was also measured before and after Adaption processing:
Measure
Before
After
Quality score
8.0
9.5
Grade
B
A
Percentile
17.8
57.7
The platform reports the quality-score change as 18.8% relative improvement. These are platform measurements of
the submitted dataset, not independent chemistry benchmarks.
The improvement is visible beyond the score. The adapted records add clearer task framing, more explicit deliverables,
better organized explanations, and stronger misconception-focused teaching while preserving the chemical problem being
solved.
BondShift training telemetry
Training curves document the run's optimization telemetry. They do not, by themselves, establish chemical correctness
or out-of-distribution generalization.
The AutoScientist-selected configuration was used unchanged.
Parameter
Value
Epochs
3
Batch size
max
Evaluations
5
Learning rate
1e-4
Scheduler
Cosine
Scheduler cycles
0.5
Warmup ratio
0.03
Minimum LR ratio
0.1
Weight decay
0
Max gradient norm
2
LoRA rank
32
LoRA alpha
64
LoRA dropout
0
Trainable modules
all-linear
Train on inputs
false
Training data
BondShift was trained from exactly 10,000 English chemistry records. The public dataset exposes the audited source
pair and the Adaption-remastered pair side by side:
Slice
Rows
BondShift mechanism core
8,017
Science bridge
1,983
Total
10,000
The final mix is approximately 20% basic, 30% intermediate, 30% advanced, and 20% expert. All six V1 chemistry
families are represented near evenly. Prompt styles range from short questions to contextual debugging requests, while
assistant responses use flexible natural prose rather than one fixed answer template.
Every row earns its place: the final 10,000 were selected from a deterministic 12,000-slot candidate plan, checked for
scope and answer-key alignment, screened for internal leakage, and deduplicated before the immutable release snapshot
was created.
What Adaption improved
The source corpus already supplied natural questions, grounded answers, and strict mechanism validation. Adaption then
added a second, enhanced view of every record:
Dataset field
Role
prompt
Original natural user question
response
Original validated answer
enhanced_prompt
Adaption-remastered instruction with clearer task framing
enhanced_completion
Adaption-remastered teaching response
reasoning_trace
Auxiliary platform-generated reasoning data
The adapted AutoScientist run is associated with the enhanced instruction/completion view. The auxiliary reasoning
trace is not BondShift's deterministic answer key and is not the archival MechanismIR sidecar.
Ground-truth architecture
The data pipeline deliberately separates five concerns:
A deterministic scenario blueprint contains only facts that may be shown to the prompt author.
A separately constructed answer key stores the expected mechanism, outcome, bond changes, misconception target,
validation targets, and provenance.
The prompt author receives the blueprint, never the hidden answer key.
The response teacher receives the frozen user prompt plus a factual grounding packet, never ideal response prose.
Validators independently check structure, leakage, restricted-template claims, and alignment with the answer key.
The training response contains the user-facing explanation and conclusion. No private chain-of-thought was generated or
exported. MechanismIR-lite is archival validation metadata and is not forced into the model's natural-language answer.
How to use
Adaption interface
Open dataset ID 9e740edb-5d61-49c0-9510-b37919676e4a in Adaption, select Interfaces, and use the BondShift
organic-chemistry mechanism companion. A strong query includes the substrate, reagents, solvent or medium, conditions,
and the exact step or conclusion that is confusing.
Local adapter inference
This release is a 1.66 GB LoRA adapter and cannot be loaded as a standalone causal language model. The repository now
contains the adapter weights, adapter configuration, tokenizer, and chat template. PEFT can read the exact base path
from adapter_config.json.
python
1import torch
2from peft import PeftConfig, PeftModel
3from transformers import AutoModelForCausalLM, AutoTokenizer
45adapter_id ="prathmeshadsod/BondShift-Llama-3.3-70B-Instruct"6peft_config = PeftConfig.from_pretrained(adapter_id)7base_model_path = peft_config.base_model_name_or_path
89tokenizer = AutoTokenizer.from_pretrained(adapter_id)10base = AutoModelForCausalLM.from_pretrained(11 base_model_path,12 torch_dtype=torch.bfloat16,13 device_map="auto",14)15model = PeftModel.from_pretrained(base, adapter_id)1617messages =[{18"role":"user",19"content":(20"Why must bromocyclohexane have an axial leaving group before an E2 "21"elimination can occur?"22),23}]24inputs = tokenizer.apply_chat_template(25 messages,26 tokenize=True,27 add_generation_prompt=True,28 return_dict=True,29 return_tensors="pt",30).to(model.device)3132with torch.inference_mode():33 output = model.generate(**inputs, max_new_tokens=384, do_sample=False)3435new_tokens = output[0][inputs["input_ids"].shape[-1]:]36print(tokenizer.decode(new_tokens, skip_special_tokens=True))
The exported configuration currently records togethercomputer/Meta-Llama-3.3-70B-Instruct-Reference. Access to a
compatible base checkpoint and hardware capable of serving a 70B model are still required. Quantization and adapter
merging should be tested separately; this repository is not a merged full checkpoint.
Intended use
BondShift is intended for:
undergraduate mechanism tutoring and formative feedback;
explaining supported arrow-pushing decisions;
debugging a student's proposed mechanism;
generating study examples within the validated V1 scope;
research on natural-language chemistry tutoring.
It is not intended to replace an instructor, validate a synthesis, provide laboratory safety instructions, or support
clinical, industrial, or high-stakes chemical decisions.
Limitations
V1 covers acid-base, resonance/formal charge, SN1, SN2, E1, and E2. Carbonyl chemistry,
electrophilic addition, rearrangement-heavy chemistry, complex aromatic substitution, radical, pericyclic,
organometallic, and broad oxidation/reduction mechanisms were not certified for this release.
The model consumes text. It was not validated as an image, sketch, molecular-graph, or SMILES interpretation system.
Stereochemical labels require complete structural and CIP information. Ambiguous prompts should receive a conditional
answer or a request for the missing structure.
Product ratios and pathway dominance can depend on concentration, temperature, solvent, and substrate detail. The V1
data intentionally avoids unsupported exact ratios, but the model can still overstate a qualitative preference.
Deterministic template consistency and automated answer-key checks reduce errors; they do not prove exhaustive chemical
correctness. Expert review remains appropriate for consequential use.
Synthetic prompt and teacher styles can transfer into the adapted model. Performance may fall on unfamiliar notation,
advanced reaction families, or adversarially incomplete questions.
The reported win rates come from the Adaption interface. No confidence intervals or public item-level evaluation set
were available for independent statistical analysis.
Only one prompt-author model and one response-teacher model were used in the accepted production release. Raw provider
attempts, validation records, deterministic provenance, repair reports, and the full metadata snapshot were retained for
auditability.
Built for the Adaption AutoScientist Challenge. Adaption provided dataset adaptation, training configuration selection,
LoRA training, and the displayed evaluation results. The production data pipeline used deterministic chemistry templates
and a large reasoning model for natural prompt and response realization, with the truth and generation paths kept
separate.