This model is a fine-tuned version of the Qwen2-1.5-Instruct using Low-Rank Adaptation (LoRA). It is specifically designed for extracting key information from bidding and bid-winning announcements. The model focuses on identifying structured data such as project names, announcement types, budget amounts, and deadlines in various formats of bidding notices.
The base model, Qwen2-1.5-Instruct, is a large-scale language model optimized for instruction-following tasks, and this fine-tuned version leverages its capabilities for precise data extraction tasks in Chinese bid announcement contexts.
Use Cases
The model can be used in applications that require the automatic extraction of structured data from text documents, particularly related to government bidding and procurement processes. For instance, based on the sample announcement, the generated output is as follows:
Fine-tuned with LoRA: The model has been adapted using LoRA, a parameter-efficient fine-tuning method, allowing it to focus on specific tasks while maintaining the power of the large base model.
Robust Information Extraction: The model is trained to extract and validate crucial fields, including budget values, submission deadlines, and industry classifications, ensuring accurate outputs even when encountering variable formats.
Language & Domain Specificity: The model excels in parsing official bidding announcements in Chinese and accurately extracting the required information for downstream processes.
Model Architecture
Base Model: Qwen2-1.5B-Instruct
Fine-Tuning Technique: LoRA
Training Data: Fine-tuned on structured and unstructured government bidding announcements
Framework: Hugging Face Transformers & PEFT (Parameter Efficient Fine Tuning)
Technical Specifications
Device Compatibility: CUDA (GPU-enabled)
Tokenization: Utilizes AutoTokenizer from Hugging Face, optimized for instruction-following tasks.
The Tongda1-1.5B-BKI model has shown remarkable performance in information extraction tasks. Compared to the baseline model Qwen2-1.5B-Instruct, Tongda1-1.5B-BKI excels across multiple evaluation metrics, particularly in extracting key information from tender announcements, achieving significant improvements. Even when compared to larger models like Qwen2.5-3B-Instruct and Qwen2-7B-Instruct, Tongda1-1.5B-BKI still demonstrates outstanding performance. Additionally, it outperforms the optimized online model glm-4-flash. Here are the evaluation results for each model:
Model
ROUGE-1
ROUGE-2
ROUGE-Lsum
BLEU
Tongda1-1.5B-BKI
0.853
0.787
0.853
0.852
Qwen2-1.5B-Instruct
0.412
0.231
0.411
0.431
Qwen2.5-3B-Instruct
0.686
0.578
0.687
0.755
Qwen2-7B-Instruct
0.703
0.578
0.703
0.789
glm-4-flash
0.774
0.655
0.775
0.816
image/png
Limitations
Language Limitation: The model is primarily trained on Chinese bidding announcements. Performance on other languages or non-bidding content may be limited.
Strict Formatting: The model may have reduced accuracy when the bidding announcements deviate significantly from common structures.
Citation
If you use this model, please consider citing it as follows:
@inproceedings{Tongda1-1.5B-BKI,
title={Tongda1-1.5B-BKI: LoRA Fine-tuned Model for Bidding Announcements},
author={Ted-Z},
year={2024}
}
Contact
For further inquiries or fine-tuning services, please contact us at Tongda.