The dataset comprises Verilog modules as entries. The entries were retrieved from the GitHub dataset on BigQuery.
For training [models (
https://huggingface.co/shailja/fine-tuned-codegen-2B-Verilog)], we filtered entries with no of characters exceeding 20000 and duplicates (exact duplicates ignoring whitespaces).
Paper: Benchmarking Large Language Models for Automated Verilog RTL Code Generation
Point of Contact: contact@shailja… See the full description on the dataset page:
https://huggingface.co/datasets/develoco/Verilog_GitHub.