Views
No views yet
Synthyra/DPLM-650M packages the airkingbd/dplm_650m checkpoint with the
FastPLMs runtime for Hugging Face Transformers. It accepts amino-acid sequences
tokenized to masked or partially masked residue IDs.trust_remote_code=True. See Technical details for each registered class and
whether its weights come from the checkpoint.1python -m pip install -r \
2 "https://huggingface.co/Synthyra/DPLM-650M/resolve/main/requirements.txt"trust_remote_code=True.1from transformers import AutoModel
2
3model_id = "Synthyra/DPLM-650M"
4model = AutoModel.from_pretrained(
5 model_id,
6 trust_remote_code=True,
7 attn_implementation="sdpa",
8).eval()model_id with the manifest-built
dist/hub/DPLM-650M path. Pass local_files_only=True.sdpa.eager, sdpa, flex_attention, flash_attention_3.
Requesting an unavailable backend raises instead of silently changing
implementation.output_attentions=True can use the documented one-call eager fallback to
materialize attention tensors. The configured backend does not change.1import torch
2from transformers import AutoTokenizer
3
4model_id = "Synthyra/DPLM-650M"
5tokenizer = AutoTokenizer.from_pretrained(
6 model_id,
7 trust_remote_code=True,
8)
9batch = tokenizer(
10 ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"],
11 padding=True,
12 return_tensors="pt",
13)
14
15with torch.inference_mode():
16 output = model(**batch)
17
18print(output.last_hidden_state.shape)1pooled = model.embed_dataset(
2 ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"],
3 batch_size=2,
4 pooling=("mean", "std"),
5)
6residues = model.embed_dataset(
7 ["MSTNPKPQRKTKRNT"],
8 full_embeddings=True,
9)
10print(pooled[0].tensor.shape) # (2 * d,)
11print(residues[0].tensor.shape) # (l, d)output and format="safetensors" or "sqlite" for transactional,
bounded-memory storage. Resume checks input order, model state, tokenizer
policy, backend, dtype, and pooling configuration before it appends data.classifier. Sequence labels have shape (b,).
Residue labels have shape (b, l) and use -100 outside biological positions.1import torch
2from transformers import AutoTokenizer
3from transformers import (
4 AutoModelForSequenceClassification,
5 AutoModelForTokenClassification,
6)
7
8model_id = "Synthyra/DPLM-650M"
9sequence_model = AutoModelForSequenceClassification.from_pretrained(
10 model_id, num_labels=2, trust_remote_code=True
11).eval()
12token_model = AutoModelForTokenClassification.from_pretrained(
13 model_id, num_labels=3, trust_remote_code=True
14).eval()
15tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
16sequences = ["MSTNPKPQRKTKRNT", "MKTIIALSYIFCLVFA"]
17batch = tokenizer(sequences, padding=True, return_tensors="pt")
18biological = batch["attention_mask"].bool()
19for special_id in tokenizer.all_special_ids:
20 biological &= batch["input_ids"].ne(special_id)
21
22sequence_labels = torch.zeros(len(sequences), dtype=torch.long)
23token_labels = torch.full_like(batch["input_ids"], -100)
24token_labels[biological] = 0
25
26with torch.inference_mode():
27 sequence_output = sequence_model(**batch, labels=sequence_labels)
28 token_output = token_model(**batch, labels=token_labels)
29print(sequence_output.logits.shape) # (b, 2)
30print(token_output.logits.shape) # (b, l, 3)python -m pip install "datasets>=4.8,<5" "peft>=0.19,<0.20"1from peft import LoraConfig, TaskType, get_peft_model
2
3peft_model = get_peft_model(
4 sequence_model,
5 LoraConfig(
6 task_type=TaskType.SEQ_CLS,
7 r=8,
8 lora_alpha=16,
9 target_modules="all-linear",
10 modules_to_save=["classifier"],
11 ),
12)classifier with the adapter.
All FastPLMs checkpoints follow the Transformers PreTrainedModel contract and
can use PEFT. The ESM2-specific shipped CLI is an example, not a
support boundary. Record the target modules, base revision, data identity, and
trainable parameter scope.1from transformers import AutoModelForMaskedLM
2
3ttt_model = AutoModelForMaskedLM.from_pretrained(
4 "Synthyra/DPLM-650M",
5 trust_remote_code=True,
6)
7metrics = ttt_model.ttt(
8 seq="MSTNPKPQRKTKRNT",
9 ttt_config={"steps": 3, "batch_size": 1, "seed": 7},
10)
11ttt_model.save_pretrained("adapted", safe_serialization=True)
12ttt_model.ttt_reset()
13print(metrics)1import torch
2from transformers import AutoModelForMaskedLM, AutoTokenizer
3
4model_id = "Synthyra/DPLM-650M"
5tokenizer = AutoTokenizer.from_pretrained(model_id)
6generator = AutoModelForMaskedLM.from_pretrained(
7 model_id,
8 trust_remote_code=True,
9).cuda().eval()
10input_ids = tokenizer("A" * 64, return_tensors="pt")["input_ids"].cuda()
11
12with torch.inference_mode():
13 generated_ids = generator.generate(input_ids, max_iter=100)
14
15sequence = tokenizer.decode(
16 generated_ids[0],
17 skip_special_tokens=True,
18).replace(" ", "")
19print(sequence)max_iter, DPLM uses the official 500-step schedule. A shorter
schedule changes the sampling process. It is not an equivalent faster mode.AutoModel omits the optional ESM pooler because this diffusion checkpoint
has no trained pooler weights. Pass add_pooling_layer=True only when you intend
to initialize and train that head.weights_license_status="resolved" and redistributable=true. Complete
publication requires all artifact, legal, parity, and atomic-publication checks.AutoConfig, AutoModel, AutoModelForMaskedLM, AutoModelForSequenceClassification, AutoModelForTokenClassificationAutoConfig = FastPLMs extension, AutoModel = pretrained, AutoModelForMaskedLM = pretrained, AutoModelForSequenceClassification = base weights + untrained task head, AutoModelForTokenClassification = base weights + untrained task headeager, sdpa, flex_attention, flash_attention_3defaultfp32_parameters_autocastrequiredcoretrueresolvedtruefalsemodels.toml. Built artifacts record exact source
identities and conversion details in source-record.json.Synthyra/DPLM-650Msource-record.jsonairkingbd/dplm_650mfastdplm_to_fastplms_v1dplmcheck, compliance, feature, artifact, benchmark0compliance tier. Its evidence identifies the
checkpoint, backend, dtype, hardware, inputs, and reference revision.apache-2.0. The local artifact contains applicable source
licenses, notices, attribution, and conversion records. Review them before use.