Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
tFINE-900m-e16-d32-1024ctx – AI Model by pszemraj | AlphaNeural AI
You can deploy this model and start earning money today!
pszemraj
/
tFINE-900m-e16-d32-1024ctx
like
0
transformers
safetensors
t5
text2text-generation
en
HuggingFaceTB/smollm-corpus
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
tFINE-900m-e16-d32-1024ctx
Pretrained T5 model with
nanoT5
:
~900m parameters, 16 layers in encoder, 32 layers in decoder
sentencepiece tokenizer with 48k vocab & byte-pair fallback
handles whitespaces etc correctly (
unlike original T5 tokenizer
)
1024 ctx during pretrain
relative_attention_num_buckets
increased to 48 from 32 for context length upscaling
Experiment logs
Training consisted of two phases:
phase one
- ~30k steps at context length 512
phase two
- 20k steps at context length 1024