Views
No views yet
| Parameter | Value |
|---|---|
| Layers | 6 |
| Attention heads | 16 |
| Embedding dimension | 512 |
| FFN hidden dimension | 2048 (GELU) |
| Vocabulary size | 10 |
| Positional encoding | Learned absolute |
| Normalization | LayerNorm (post-residual, eps=1e-12) |
| Architecture | Post-LN BERT encoder |
| Max sequence length | 510 nt (512 tokens; fixed-length training) |
| Checkpoint size | ~19.5M parameters |
[PAD]=0, [UNK]=1, [CLS]=2, [SEP]=3,
[MASK]=4, N=5, A=6, C=7, G=8, T=9. Input U is
normalized to T.SpliceBERT.510nt/pytorch_model.bin (from zenodo:7995778)eager and sdpa attention backends.
Verified on GPU with PyTorch 2.7.1 / CUDA 12.9.| Model | Context | Training data | Notes |
|---|---|---|---|
| SpliceBERT-1024nt | 1024 nt | 72 vertebrates | Variable-length; general purpose |
| SpliceBERT-510nt | 510 nt (fixed) | 72 vertebrates | This model |
| SpliceBERT-human-510nt | 510 nt (fixed) | Human only | Human-specific |
1import torch
2from transformers import AutoTokenizer, AutoModel
3
4tokenizer = AutoTokenizer.from_pretrained("Taykhoom/SpliceBERT-510nt", trust_remote_code=True)
5model = AutoModel.from_pretrained("Taykhoom/SpliceBERT-510nt", trust_remote_code=True)
6model.eval()
7
8# The model was trained on exactly 510 nt; tokenizer handles U->T automatically
9seq = ("ATCGATCG" * 64)[:510] # exactly 510 nt
10enc = tokenizer(seq, return_tensors="pt")
11
12with torch.no_grad():
13 out = model(**enc, output_hidden_states=True)
14
15hidden = out.last_hidden_state[0] # (512, 512)
16token_emb = hidden[1:-1] # strip [CLS] and [SEP] -> (510, 512)
17mean_emb = token_emb.mean(dim=0) # (512,)BertForMaskedLM with transformers==4.20.1.
This port uses BERT-updated, which
adds attn_implementation="sdpa" and attn_implementation="flash_attention_2" support
not present in the original codebase.pooler.dense) are not present in the original checkpoint and are
not included in the saved model.safetensors. add_pooling_layer=True (the default)
allocates the pooler layer but its weights are randomly initialized -- do not use
pooler_output without fine-tuning.1@article{chen2024_splicebert,
2 title = {Self-supervised learning on millions of primary {RNA} sequences from 72 vertebrates improves sequence-based {RNA} splicing prediction},
3 author = {Chen, Ken and Zhou, Yue and Ding, Maolin and Wang, Yu and Ren, Zhixiang and Yang, Yuedong},
4 journal = {Briefings in Bioinformatics},
5 volume = {25},
6 number = {3},
7 pages = {bbae163},
8 year = {2024},
9 doi = {10.1093/bib/bbae163}
10}