The first model in the Wongwian series. A compact 272M-parameter Thai-centric language model trained entirely from scratch on 40B tokens of Thai-dominant data and then instruction-fine-tuned. No base weights from existing open models were used at any stage.
"Cultural & Localization AI for the world" — Wongwian
Wongwian Micro Instruct — 272M (ภาษาไทย)
โมเดลตัวแรกในซีรีส์ Wongwian เป็นโมเดลภาษาขนาดเล็ก 272 ล้านพารามิเตอร์ที่เน้นภาษาไทยเป็นหลัก ฝึกขึ้นมาจากศูนย์ (train from scratch) บนข้อมูล 40 พันล้าน token แล้วผ่านการ instruction fine-tuning โดยไม่ได้นำ open weights ของโมเดลอื่นมาต่อยอดแต่อย่างใด
"Cultural & Localization AI for the world" — Wongwian
ภาพรวมโมเดล
คุณสมบัติ
ค่า
ตระกูล
Wongwian
ซีรีส์
Micro
เวอร์ชัน
Instruct v1 (step 450)
สถาปัตยกรรม
LlamaForCausalLM
จำนวนพารามิเตอร์
~272 ล้าน
Context length
2,048 tokens
Precision
bfloat16
Token ที่ใช้ pre-train
~40 พันล้าน
ภาษาหลัก
ไทย 🇹🇭
ภาษารอง
อังกฤษ 🇬🇧
License
MIT
เกี่ยวกับ Wongwian
Wongwian คือโครงการวิจัย AI ภาษาไทยที่มุ่งสร้าง โมเดลภาษา AI ที่มีประสิทธิภาพสูงและเข้าใจบริบทเชิงวัฒนธรรม สำหรับภาษาไทยและภาษาอื่น ๆ ที่ขาดแคลนทรัพยากร โครงการนี้แสดงให้เห็นว่าโมเดล AI ที่มีคุณภาพสูงไม่จำเป็นต้องมีพารามิเตอร์มหาศาล โมเดลขนาดเล็กที่ออกแบบอย่างพิถีพิถันและใช้ข้อมูลที่เหมาะสม สามารถตอบโจทย์การใช้งานจริงได้อย่างมีประสิทธิผล
วิสัยทัศน์:Cultural & Localization AI for the world — สร้าง AI ที่เข้าใจความละเอียดอ่อนทางภาษา วัฒนธรรม และบริบทของแต่ละชุมชนอย่างลึกซึ้ง
Wongwian is a Thai AI research initiative focused on building efficient, culturally-grounded language models for Thai and other under-resourced languages. The project demonstrates that high-quality language AI does not require massive parameter counts — a carefully designed small model, trained on the right data, can serve real-world use cases effectively.
Vision:Cultural & Localization AI for the world — delivering language models that deeply understand the cultural, linguistic, and contextual nuances of each community they serve.
What makes this model different:
Trained from scratch on Thai-dominant data — not a fine-tune of any existing public model
Custom SentencePiece Unigram tokenizer (32K vocab) designed for Thai morphology and mixed Thai-English text
Size: At 272M parameters, this model is suited for assistive tasks and demonstrations, not advanced reasoning or complex long-form analysis.
Hallucination: Like all language models, the model may produce inaccurate information. Always verify critical outputs.
Context length: Maximum 2 048 tokens per call; performance may degrade near the limit.
Bias: The model may reflect biases present in the training corpus.
Citation
bibtex
1@misc{wongwian2026micro,
2 title = {Wongwian Micro Instruct: A Thai-Centric Small Language Model Trained from Scratch},
3 author = {Wongwian AI Research},
4 year = {2026},
5 url = {https://huggingface.co/wongwian-org/wongwian-micro-instruct}
6}