RWKV-4 430M is a L24-D1024 causal language model trained on the Pile. See
https://github.com/BlinkDL/RWKV-LM for details.
Final checkpoint:
RWKV-4-Pile-430M-20220808-8066.pth : Trained on the Pile for 333B tokens.
With tiny attention (--tiny_att_dim 512 --tiny_att_layer 18):
RWKV-4a-Pile-433M-20221223-8039.pth