Komdigi-3b-pad-merged
Model Description
Komdigi-3b-pad-merged is a unified multi-task Vision-Language Model developed for the Perlindungan Anak di Ruang Digital (PAD) use case. Unlike a model that is limited to image analysis, this model is designed to handle multiple PAD-related tasks within a single architecture.
The model is trained as a unified VLM to perform:
- Text classification with reasoning
- Image classification with reasoning
- Keyword generation
The main purpose of the model is to support content analysis, age-rating classification, reasoning generation, and keyword extraction in the context of digital child protection. It is intended for research, prototyping, and human-in-the-loop moderation systems.
Model Details
| Attribute | Value |
|---|
| Model name | aitf-komdigi/Komdigi-3b-pad-merged |
| Model type | Unified Vision-Language Model |
| Architecture | Mistral3 / Pixtral-style Vision-Language Model |
| Base model | aitf-komdigi/KomdigiITS-3B-PAD-CPT |
| Parameter size | Approximately 4B parameters |
| Tensor type | BF16 |
| File format | Safetensors |
| Main language | Indonesian |
| Secondary language | English |
| Domain | Perlindungan Anak di Ruang Digital |
| Training approach | Multi-task supervised fine-tuning |
| Model status | Merged model |
| Supported tasks | Text classification, image classification, keyword generation |
Supported Tasks
1. Text Classification with Reasoning
The model can analyze text-based content such as captions, posts, comments, or short-form digital content. It predicts the appropriate age-rating or content-safety category and provides a reasoning explanation for the prediction.
Example use cases:
- analyzing social media captions,
- classifying text-based digital content,
- generating explanation for the assigned rating,
- supporting human moderation review.
2. Image Classification with Reasoning
The model can analyze visual content and classify it into the appropriate PAD rating category. It can also generate reasoning based on visual evidence observed in the image.
Example use cases:
- analyzing social media images,
- detecting content that may not be suitable for certain age groups,
- explaining visual factors that influence the classification,
- supporting review of image-based content.
3. Keyword Generator
The model can generate relevant keywords from a given content context. This task is useful for indexing, search, moderation tagging, and content understanding.
Example use cases:
- generating moderation keywords,
- extracting relevant content tags,
- supporting content retrieval,
- improving downstream filtering or recommendation systems.
Intended Use
Direct Use
This model is intended for PAD-related content analysis, including:
- classifying text content with reasoning,
- classifying image content with reasoning,
- generating relevant keywords from content context,
- supporting content moderation workflows,
- assisting age-rating recommendation,
- supporting Indonesian child-safety research and prototyping.
Downstream Use
The model may be integrated into larger systems such as:
- content moderation dashboards,
- child-safety screening tools,
- age-rating recommendation systems,
- keyword tagging systems,
- content indexing pipelines,
- human-in-the-loop moderation platforms.
Out-of-Scope Use
This model should not be used as the sole authority for:
- final legal or regulatory decisions,
- fully automated content takedown,
- law enforcement decisions,
- identifying or profiling individuals,
- determining the real age or identity of a person,
- surveillance or invasive monitoring,
- medical, psychological, or legal assessment.
The model should assist human reviewers, not replace human judgment.
Rating Categories
The model is trained to classify content into the following PAD rating categories:
| Rating | General Meaning |
|---|
Semua Umur | Content that is generally safe, neutral, educational, or child-friendly. |
7+ | Content that may require light age guidance for children above 7 years old. |
13+ | Content that may contain themes more suitable for teenagers. |
15+ | Content that may contain stronger themes requiring older teen maturity. |
18+ | Content suitable only for adults. |
Konten Terlarang | Content that violates safety rules or should not be distributed. |
Unrated | Content that cannot be confidently classified or does not fit the available categories. |
The exact interpretation of each category should follow the PAD annotation guideline used during dataset construction.
Recommended Output Format
For classification tasks, the model can be prompted to return structured JSON.
1{
2 "rating": "string",
3 "reason": "string"
4}
For keyword generation, the recommended output format is:
1{
2 "keywords": ["keyword_1", "keyword_2", "keyword_3"]
3}
For a combined response, the following schema can be used:
1{
2 "task": "classification_or_keyword_generation",
3 "rating": "string",
4 "reason": "string",
5 "keywords": ["keyword_1", "keyword_2", "keyword_3"]
6}
Training Details
Training Dataset Composition
The model was trained using a unified multi-task supervised fine-tuning strategy. The training data combines task-specific datasets that were converted into a shared instruction-following format, allowing the same VLM architecture to learn text classification, image classification, reasoning generation, and keyword generation in a single training workflow.
The training notebook uses a combined multi-task dataset loaded from:
nuresens/PAD-Combined-Dataset_v6
The dataset is separated into three task sources using the source field:
| Source | Task | Number of Samples |
|---|
pad1 | Text Classification with Reasoning | 13,406 |
pad2 | Keyword Generator | 15,030 |
pad3 | Image Classification with Reasoning | 62,427 |
| Total | All Tasks | 90,863 |
The three sources are merged into a single metadata dataset containing the following columns:
| Column | Description |
|---|
messages | Chat-style instruction data containing user prompts and assistant responses. |
source | Dataset source identifier: pad1, pad2, or pad3. |
source_idx | Original row index from each source dataset. |
has_image | Boolean marker indicating whether the sample contains an image placeholder. |
Feature Structure
| Column | Type | Description |
|---|
image | Image | Visual input used by the vision-language model for image-based classification and reasoning. |
messages | List | Conversation-style instruction data containing user prompts and assistant responses. |
The messages field follows a chat-style structure consisting of role and content. This structure allows the dataset to support multimodal supervised fine-tuning across different task instructions. Text classification and keyword generation samples are text-only instruction samples, while image classification samples include image placeholders and are lazily paired with their original images during collation.
Dataset Statistics
1. Image Classification and Reasoning Dataset
The image classification and reasoning dataset consists of 62,427 samples. The dataset is used to train the model to analyze visual content and predict the appropriate PAD rating category with an accompanying explanation.
| Rating | Number of Samples | Percentage |
|---|
| Semua Umur | 15,973 | 25.59% |
| 13+ | 9,422 | 15.09% |
| 7+ | 8,189 | 13.12% |
| Konten Terlarang | 7,676 | 12.30% |
| Unrated | 7,497 | 12.01% |
| 18+ | 6,851 | 10.97% |
| 15+ | 6,819 | 10.92% |
| Total | 62,427 | 100.00% |
The largest category in the image classification dataset is Semua Umur, representing 25.59% of the dataset. The remaining categories are distributed between approximately 10.92% and 15.09%, indicating a moderately imbalanced dataset.
2. Text Classification and Reasoning Dataset
The text classification and reasoning dataset consists of 13,406 rows. This dataset is used to train the model to classify text-based content into age-rating categories and generate explanations that justify the classification.
| Label Rating Usia | Main Content Characteristics | Number of Rows | Percentage |
|---|
| Konten Terlarang | Judi online, kekerasan verbal, pelecehan, SARA | 4,497 | 33.54% |
| Semua Umur | Ramah keluarga, aktivitas keseharian yang positif | 3,252 | 24.26% |
| 15+ | Diskusi sosial-politik, kriminalitas tanpa glorifikasi | 1,826 | 13.62% |
| 13+ | Curahan hati remaja, konflik sosial pertemanan | 1,458 | 10.88% |
| 7+ | Edukasi ringan, hiburan anak, persaingan olahraga | 1,239 | 9.24% |
| 18+ | Romansa dewasa, gaya hidup malam, edukasi self-harm | 1,134 | 8.46% |
| Total | Dataset Master SFT | 13,406 | 100.00% |
The text classification dataset is more imbalanced than the image classification dataset. The largest category is Konten Terlarang, representing 33.54% of the dataset, followed by Semua Umur at 24.26%. The smaller categories, especially 18+ and 7+, should be monitored carefully during evaluation using macro-averaged metrics.
3. Keyword Generator Dataset
The keyword generator dataset consists of 15,030 records. This dataset is used to train the model to generate a variable number of relevant keywords from a given content context.
| Number of Keywords | Number of Records | Percentage |
|---|
| 1 | 845 | 5.62% |
| 2 | 2,347 | 15.62% |
| 3 | 4,700 | 31.27% |
| 4 | 3,691 | 24.56% |
| 5 | 1,528 | 10.17% |
| 6 | 671 | 4.46% |
| 7 | 327 | 2.18% |
| 8 | 238 | 1.58% |
| 9 | 218 | 1.45% |
| 10 | 275 | 1.83% |
| 11 | 65 | 0.43% |
| 12 | 57 | 0.38% |
| 13 | 32 | 0.21% |
| 14 | 23 | 0.15% |
| 15 | 10 | 0.07% |
| 16 | 1 | 0.01% |
| Total | 15,030 | 100.00% |
The keyword generator dataset is concentrated around outputs containing 3 to 4 keywords, which together account for 55.83% of all records. The average number of keywords per record is approximately 3.81, with a median and mode of 3 keywords. This indicates that the keyword generation task is primarily optimized for concise keyword outputs rather than long keyword lists.
Dataset Split
The dataset is split using a source-aware strategy. Each source is split independently using an 80:10:10 ratio, then recombined into train, validation, and test splits. This ensures that all tasks remain represented proportionally in each split.
Overall Split
| Split | Number of Samples | Percentage |
|---|
| Train | 72,689 | 80.00% |
| Validation | 9,087 | 10.00% |
| Test | 9,087 | 10.00% |
| Total | 90,863 | 100.00% |
Split by Task Source
| Source | Task | Train | Validation | Test | Total |
|---|
pad1 | Text Classification with Reasoning | 10,724 | 1,341 | 1,341 | 13,406 |
pad2 | Keyword Generator | 12,024 | 1,503 | 1,503 | 15,030 |
pad3 | Image Classification with Reasoning | 49,941 | 6,243 | 6,243 | 62,427 |
| Total | All Tasks | 72,689 | 9,087 | 9,087 | 90,863 |
Training Procedure
This model was trained using a unified multi-task fine-tuning strategy. Instead of training separate models for text classification, image classification, and keyword generation, all tasks were learned by a single Vision-Language Model.
The model was fine-tuned end-to-end using task-specific prompts and responses in a shared multimodal instruction format. This allows the model to preserve a unified latent representation across tasks and reduces the risk of performance degradation caused by separate adapter merging.
| Component | Value |
|---|
| Base model | aitf-komdigi/KomdigiITS-3B-PAD-CPT |
| Model architecture | Unified Vision-Language Model |
| Training approach | Multi-task supervised fine-tuning |
| Tasks | Text classification, image classification, keyword generation |
| Data format | TRL-style messages |
| Split strategy | Source-aware 80:10:10 train/validation/test split |
| Merge status | Merged model |
| Precision | BF16 |
| Frameworks | Transformers, TRL, PEFT, Unsloth |
Training Configuration
The model was fine-tuned using Unsloth FastVisionModel and TRL SFTTrainer. Training was performed as multi-task supervised fine-tuning over the merged dataset containing text classification, image classification, and keyword generation tasks.
Base Model and Sequence Configuration
| Component | Value |
|---|
| Base model | aitf-komdigi/KomdigiITS-3B-PAD-CPT |
| Model class | FastVisionModel |
| Maximum sequence length | 4096 |
| Load in 4-bit | False |
| Precision | BF16 if supported, otherwise FP16 |
| Gradient checkpointing | unsloth |
| Chat template reference | mistralai/Ministral-3-3B-Instruct-2512 |
| Padding side | Right padding |
The notebook applies the chat template from mistralai/Ministral-3-3B-Instruct-2512 to the model processor and tokenizer so that the training format follows the instruction-style format expected by the model.
LoRA Configuration
The model was fine-tuned using LoRA. Both the vision and language components were enabled for fine-tuning.
| Parameter | Value |
|---|
LoRA rank r | 16 |
| LoRA alpha | 16 |
| LoRA dropout | 0.0 |
| Bias | none |
| Target modules | all-linear |
| Random state | 42 |
| Fine-tune vision layers | True |
| Fine-tune language layers | True |
| Fine-tune attention modules | True |
| Fine-tune MLP modules | True |
The training run reported 33,751,040 trainable parameters out of 3,882,841,088 total parameters, meaning approximately 0.87% of the model parameters were trained.
Data Collation
The notebook uses UnslothVisionDataCollator with a custom lazy image collation strategy. Images are loaded only when a sample has an image marker, which reduces unnecessary image decoding for text-only samples.
| Component | Value |
|---|
| Base collator | UnslothVisionDataCollator |
| Custom wrapper | Lazy vision collator |
| Train on responses only | True |
| Instruction part | [INST] |
| Response part | [/INST] |
| Image loading strategy | Lazy image injection for pad3 samples only |
The training loss is computed only on the assistant response, which is useful for instruction tuning because the model is optimized to generate the expected output rather than reproduce the full prompt.
SFT Training Arguments
| Hyperparameter | Value |
|---|
| Per-device train batch size | 4 |
| Gradient accumulation steps | 2 |
| Effective train batch size | 8 |
| Per-device evaluation batch size | 4 |
| Number of epochs | 2 |
| Learning rate | 2e-4 |
| Warmup ratio | 0.03 |
| Learning rate scheduler | Cosine |
| Weight decay | 0.01 |
| Optimizer | adamw_8bit |
| Max gradient norm | 1.0 |
| Logging steps | 10 |
| Evaluation strategy | Steps |
| Evaluation steps | 1000 |
| Save strategy | Steps |
| Save steps | 1000 |
| Save total limit | 2 |
| Seed | 42 |
| Output directory | outputs_pad_sft |
| Report to | Weights & Biases |
| Dataset text field | Empty string |
| Skip dataset preparation | True |
| Dataset number of processes | 4 |
| Dataloader workers | 2 |
| Pin memory | True |
Training Run Summary
| Attribute | Value |
|---|
| Number of training examples | 72,689 |
| Number of epochs | 2 |
| Total optimization steps | 18,174 |
| Number of GPUs | 1 |
| GPU used | NVIDIA A100-SXM4-40GB |
| Total batch size | 8 |
| Trainable parameters | 33,751,040 |
| Total parameters | 3,882,841,088 |
| Percentage of trained parameters | 0.87% |
Evaluation
The model was evaluated across three main tasks: text classification with reasoning, image classification with reasoning, and keyword generation. The evaluation results show that the unified model is able to maintain strong performance across tasks without significant signs of negative interference.
Text Classification with Reasoning
| Metric | Score |
|---|
| Accuracy | 0.9276 |
| Macro F1 | 0.9129 |
| Weighted F1 | 0.9267 |
| ROUGE-1 | 0.5316 |
| ROUGE-2 | 0.2743 |
| ROUGE-L | 0.3999 |
| BERTScore Precision | 0.8136 |
| BERTScore Recall | 0.8121 |
| BERTScore F1 | 0.8126 |
Image Classification with Reasoning
| Metric | Score |
|---|
| Accuracy | 0.9121 |
| Macro F1 | 0.9153 |
| Weighted F1 | 0.9120 |
| ROUGE-1 | 0.4990 |
| ROUGE-2 | 0.2360 |
| ROUGE-L | 0.3793 |
| BERTScore Precision | 0.8023 |
| BERTScore Recall | 0.8004 |
| BERTScore F1 | 0.8012 |
Keyword Generator
| Metric | Score |
|---|
| Accuracy | 0.9261 |
| ROUGE-L Precision | 0.7216 |
| ROUGE-L Recall | 0.7302 |
| ROUGE-L F1-Score | 0.7252 |
Limitations
The model has several limitations:
- The model may produce incorrect predictions for ambiguous or context-dependent content.
- The model may be sensitive to prompt wording and output-format instructions.
- For text classification, the model may misinterpret slang, sarcasm, coded language, or culturally specific expressions.
- For image classification, the model may under-detect subtle harmful content when visual evidence is unclear.
- For keyword generation, the model may generate overly broad or incomplete keywords if the input context is vague.
- The model may generate explanations that sound plausible but are not fully grounded in the input.
- The model may not fully capture legal, cultural, or platform-specific policy nuances.
- The model should not be used without human review for high-impact moderation decisions.
Bias, Risks, and Safety Considerations
Content moderation and age-rating models can reflect biases in their datasets, annotation guidelines, and policy definitions. Performance may vary across language, cultural context, visual style, demographic representation, and content source.
Potential risks include:
- false positives that incorrectly flag safe content,
- false negatives that miss harmful or restricted content,
- inconsistent treatment of culturally specific content,
- unsupported or hallucinated explanations,
- over-reliance on automated moderation,
- irrelevant or misleading generated keywords.
Recommended mitigations:
- Use human-in-the-loop review for sensitive cases.
- Evaluate the model separately for text, image, and keyword generation tasks.
- Monitor per-class performance, especially for minority rating categories.
- Use strict JSON schemas during inference.
- Keep audit logs for model predictions in production systems.
- Re-evaluate the model after dataset, taxonomy, or policy changes.
- Combine the model with rule-based policy checks and escalation mechanisms.
Recommended Deployment Setting
For deployment, the model should be used as part of a larger PAD analysis pipeline:
- Content is submitted as text, image, or image-text input.
- The system routes the input to the appropriate task prompt.
- The model predicts the rating, generates reasoning, or produces keywords.
- A validation layer checks JSON format and allowed labels.
- High-risk or uncertain cases are escalated to human reviewers.
- Final decisions and model outputs are logged for audit and future evaluation.
Ethical Considerations
This model is developed to support safer digital content access and child protection. It should be used responsibly, proportionally, and transparently. The model should not be used for surveillance, profiling, punitive action, or automated enforcement without appropriate human oversight, governance, and accountability.
Citation
If you use this model, please cite it as:
1@misc{komdigi_3b_pad_merged,
2 title = {Komdigi-3b-pad-merged: A Unified Vision-Language Model for Text Classification, Image Classification, and Keyword Generation in Perlindungan Anak di Ruang Digital},
3 author = {AITF Komdigi},
4 year = {2026},
5 publisher = {Hugging Face},
6 howpublished = {\url{https://huggingface.co/aitf-komdigi/Komdigi-3b-pad-merged}}
7}
Contact
For questions, issues, or collaboration, please use the Hugging Face repository discussion page.