Views
No views yet
Parser Network from the StructFormer. It was trained for the BabyLM 2024 challenge's Strict-Small track.| Hyperparameter | Value |
|---|---|
| Initial learning rate | 5e-3 |
| Batch size | 256 |
| Steps | 13495 |
| shuffled | True |
| attention_probs_dropout_prob | 0.1 |
| classifier_dropout | 0.2 |
| hidden_dropout_prob | 0.1 |
| hidden_size | 384 |
| intermediate_size | 1024 |
| layer_norm_eps | 1e-07 |
| max_position_embeddings | 512 |
| num_attention_heads | 6 |
| num_hidden_layers | 12 |
| vocab_size | 16384 |
| n_parser_layers | 4 |
| parser_conv_size | 9 |
| Hyperparameter | Value |
|---|---|
| Initial learning rate | 5e-5 |
| Batch size | 64 |
| Maximum epochs | 10 |
| Evaluate every (epochs) | 1 |
| Patience | 10 (for CoLA, MRPC, RTE, BoolQ, MultiRC, and WSC), 100 (for MNLI, MNLI-MM, QQP, QNLI, and SST-2) |
| Seed | 12 |
1@misc{shen2020structformer,
2 title={StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language Modeling},
3 author={Yikang Shen and Yi Tay and Che Zheng and Dara Bahri and Donald Metzler and Aaron Courville},
4 year={2020},
5 eprint={2012.00857},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL}}```1@inproceedings{georges-gabriel-charpentier-samuel-2023-layers,
2title = "Not all layers are equally as important: Every Layer Counts {BERT}",
3author = "Georges Gabriel Charpentier, Lucas and
4 Samuel, David",
5editor = "Warstadt, Alex and
6 Mueller, Aaron and
7 Choshen, Leshem and
8 Wilcox, Ethan and
9 Zhuang, Chengxu and
10 Ciro, Juan and
11 Mosquera, Rafael and
12 Paranjabe, Bhargavi and
13 Williams, Adina and
14 Linzen, Tal and
15 Cotterell, Ryan",
16booktitle = "Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning",
17month = dec,
18year = "2023",
19address = "Singapore",
20publisher = "Association for Computational Linguistics",
21url = "https://aclanthology.org/2023.conll-babylm.20",
22doi = "10.18653/v1/2023.conll-babylm.20",
23pages = "238--252",
24}```