New embeddings were initialized by averaging existing sub-token embeddings.
The resulting embedding was inserted into the expanded embedding matrix.
-
Original embeddings unchanged:
True
-
New vocabulary size:
30537 tokenizer tokens
-
Embedding matrix:
30592 x 768
-
End-to-end inference verified successfully.