The method is the same as the stella-v2, I just extend the length of the context on tao.(I found if you want to use the fully-8k context, you maybe need to convert the model to float32).
Now I'm working on the tao-v2, It will have a different sturcture.