Views
No views yet
Synthyra/ESMplusplus_small packages the biohub/ESMC-300M checkpoint with
the FastPLMs runtime for Hugging Face Transformers. It accepts amino-acid
sequences tokenized to residue IDs.trust_remote_code=True. See Technical details for each registered class and
whether its weights come from the checkpoint.1python -m pip install -r \
2 "https://huggingface.co/Synthyra/ESMplusplus_small/resolve/main/requirements.txt"trust_remote_code=True.1from transformers import AutoModel
2
3model_id = "Synthyra/ESMplusplus_small"
4model = AutoModel.from_pretrained(
5 model_id,
6 trust_remote_code=True,
7 attn_implementation="sdpa",
8).eval()model_id with the manifest-built
dist/hub/ESMplusplus_small path. Pass local_files_only=True.sdpa.eager, sdpa, flex_attention, flash_attention_2,
flash_attention_3. Requesting an unavailable backend raises instead of
silently changing implementation.output_attentions=True can use the documented one-call eager fallback to
materialize attention tensors. The configured backend does not change.1import torch
2from transformers import AutoTokenizer
3
4model_id = "Synthyra/ESMplusplus_small"
5tokenizer = AutoTokenizer.from_pretrained(
6 model_id,
7 trust_remote_code=True,
8)
9batch = tokenizer(
10 ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"],
11 padding=True,
12 return_tensors="pt",
13)
14
15with torch.inference_mode():
16 output = model(**batch)
17
18print(output.last_hidden_state.shape)1pooled = model.embed_dataset(
2 ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"],
3 batch_size=2,
4 pooling=("mean", "std"),
5)
6residues = model.embed_dataset(
7 ["MSTNPKPQRKTKRNT"],
8 full_embeddings=True,
9)
10print(pooled[0].tensor.shape) # (2 * d,)
11print(residues[0].tensor.shape) # (l, d)output and format="safetensors" or "sqlite" for transactional,
bounded-memory storage. Resume checks input order, model state, tokenizer
policy, backend, dtype, and pooling configuration before it appends data.classifier. Sequence labels have shape (b,).
Residue labels have shape (b, l) and use -100 outside biological positions.1import torch
2from transformers import AutoTokenizer
3from transformers import (
4 AutoModelForSequenceClassification,
5 AutoModelForTokenClassification,
6)
7
8model_id = "Synthyra/ESMplusplus_small"
9sequence_model = AutoModelForSequenceClassification.from_pretrained(
10 model_id, num_labels=2, trust_remote_code=True
11).eval()
12token_model = AutoModelForTokenClassification.from_pretrained(
13 model_id, num_labels=3, trust_remote_code=True
14).eval()
15tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
16sequences = ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"]
17batch = tokenizer(sequences, padding=True, return_tensors="pt")
18biological = batch["attention_mask"].bool()
19for special_id in tokenizer.all_special_ids:
20 biological &= batch["input_ids"].ne(special_id)
21
22sequence_labels = torch.zeros(len(sequences), dtype=torch.long)
23token_labels = torch.full_like(batch["input_ids"], -100)
24token_labels[biological] = 0
25
26with torch.inference_mode():
27 sequence_output = sequence_model(**batch, labels=sequence_labels)
28 token_output = token_model(**batch, labels=token_labels)
29print(sequence_output.logits.shape) # (b, 2)
30print(token_output.logits.shape) # (b, l, 3)python -m pip install "datasets>=4.8,<5" "peft>=0.19,<0.20"1from peft import LoraConfig, TaskType, get_peft_model
2
3peft_model = get_peft_model(
4 sequence_model,
5 LoraConfig(
6 task_type=TaskType.SEQ_CLS,
7 r=8,
8 lora_alpha=16,
9 target_modules="all-linear",
10 modules_to_save=["classifier"],
11 ),
12)classifier with the adapter.
All FastPLMs checkpoints follow the Transformers PreTrainedModel contract and
can use PEFT. The ESM2-specific shipped CLI is an example, not a
support boundary. Record the target modules, base revision, data identity, and
trainable parameter scope.1from transformers import AutoModelForMaskedLM
2
3ttt_model = AutoModelForMaskedLM.from_pretrained(
4 "Synthyra/ESMplusplus_small",
5 trust_remote_code=True,
6)
7metrics = ttt_model.ttt(
8 seq="MSTNPKPQRKTKRNT",
9 ttt_config={"steps": 3, "batch_size": 1, "seed": 7},
10)
11ttt_model.save_pretrained("adapted", safe_serialization=True)
12ttt_model.ttt_reset()
13print(metrics)sequence_id is supplied, it controls ESMC attention groups and padding.
attention_mask is ignored. Values greater than or equal to zero are valid
sequence-group IDs. -1 marks padding. Omit sequence_id to use
attention_mask for padding.1import torch
2
3model.load_sae_models("biohub/ESMC-300M-sae-layer23-k64-codebook65536", [23])
4
5with torch.inference_mode():
6 output = model(**batch, normalize_sae=True)
7
8features = output.sae_outputs["layer23"]
9print(features.shape, features.layout) # (valid_token_count, codebook_dim), sparse COOload_sae_models reads the shared config.json and one
layer_{index}.safetensors shard per requested layer, from a Hub repository
or a local directory, and attaches the layers on the model device in the model
dtype. add_sae_models still accepts official Biohub ESMCSAEModel.layers
entries.compute_sae=False to skip SAE work.
Outputs are detached sparse tensors with keys such as layer{N}. They omit
padding. The model uses sequence_id, then attention_mask, for padding.
normalize_sae=True uses Biohub (features / max) * idf normalization. SAE
computation requires input_ids. It rejects mask tokens because Biohub trained
the SAEs with unmasked sequences. This interface supports hidden-state SAEs
only, not MLP-output SAEs. FastPLMs does not copy SAE weights or add SAE
checkpoints to its model manifest.1import torch
2from transformers import AutoModel
3
4fp8_model = AutoModel.from_pretrained(
5 "Synthyra/ESMplusplus_small",
6 trust_remote_code=True,
7 dtype=torch.bfloat16,
8).cuda().eval()
9fp8_model.enable_fp8()
10print(fp8_model.esmc_precision_status)
11
12with torch.inference_mode():
13 fp8_output = fp8_model(**{name: value.cuda() for name, value in batch.items()})torch.inference_mode(). The model pads the sequence
dimension to a multiple of 16. Transformer Engine converts supported linear
layers. The call fails if the dependency, compatible CUDA hardware, or complete
conversion set is unavailable. It does not silently use BF16. FP8 does not
claim numerical parity.| Backend | Support | Measurement status |
|---|---|---|
sdpa | Recommended fidelity path | Pending release measurement |
eager | Supported | Pending release measurement |
flash_attention_2 | Supported | Unavailable on current GH200/aarch64 lock |
flex_attention | Supported, numerically divergent | Pending release measurement |
flash_attention_3 | Supported, numerically divergent | Unavailable on current GH200/aarch64 lock |
AutoConfig, AutoModel, AutoModelForMaskedLM, AutoModelForSequenceClassification, AutoModelForTokenClassificationAutoConfig = FastPLMs extension, AutoModel = pretrained, AutoModelForMaskedLM = pretrained, AutoModelForSequenceClassification = base weights + untrained task head, AutoModelForTokenClassification = base weights + untrained task headeager, sdpa, flex_attention, flash_attention_2, flash_attention_3default, fp8 (experimental)static_parametersnot_applicablecoretrueresolvedtruefalsemodels.toml. Built artifacts record exact source
identities and conversion details in source-record.json.Synthyra/ESMplusplus_smallsource-record.jsonbiohub/ESMC-300Mfastesmc_to_fastplms_v1biohub-esm, biohub-transformerscheck, compliance, feature, artifact, benchmark0compliance tier. Its evidence identifies the
checkpoint, backend, dtype, hardware, inputs, and reference revision.mit. The local artifact contains applicable source
licenses, notices, attribution, and conversion records. Review them before use.