Model Card for Model ID
This model is a fine-tuned version of the meta-llama/Llama-3.1-8B-Instruct model, specialized in classifying textual content from dark web sources into a range of illicit and non-illicit content categories. The model was trained using QLoRA and PEFT for efficient low-resource finetuning.
Model Details
Model Description
This model was developed to identify and classify various dark web content categories, such as Cryptocurrency, Drugs_Illegal, Porno_Child-pornography, Hosting_Software, and many others. The training prompts follow a few-shot format that helps the LLM understand the classification context clearly.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
- Developed by: [Quoc Khoa Tran]
- Funded by [optional]: [Monash University]
- Shared by [optional]: [More Information Needed]
- Model type: [Causal Language Model (LLM for classification via instruction following)]
- Language(s) (NLP): [Multi-languages]
- License: [Apache 2.0 (inherited from base model)]
- Finetuned from model [optional]: [meta-llama/Llama-3.1-8B-Instruct]
Model Sources [optional]
- Repository: [More Information Needed]
- Paper [optional]: [More Information Needed]
- Demo [optional]: [More Information Needed]
Uses
Direct Use
This model is designed to be used as a text classification tool to label scraped or user-submitted content from dark web sources into pre-defined categories.
[More Information Needed]
Downstream Use [optional]
Integration into security monitoring systems
Filtering or tagging content for moderation
Academic research on illicit activity classification
[More Information Needed]
Out-of-Scope Use
Real-time moderation of critical systems without human oversight
Classifying general-purpose content outside of dark web contexts
[More Information Needed]
Bias, Risks, and Limitations
The model may incorrectly classify borderline or ambiguous content.
Imbalanced categories in training may bias predictions toward certain classes.
Misuse may lead to false positives in sensitive applications (e.g., legal enforcement).
[More Information Needed]
Recommendations
Always use human review in high-risk scenarios.
Evaluate the model in your own context before deployment.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
How to Get Started with the Model
Use the code below to get started with the model.
[More Information Needed]
Training Details
Training Data
The dataset consists of 4,178 samples of scraped text from dark web pages, each labeled with one of 40+ content categories (e.g., Hosting_File-sharing, Violence_Weapons, Social-Network_Blog).
[More Information Needed]
Training Procedure
Preprocessing [optional]
[More Information Needed]
Training Hyperparameters
- Training regime: [More Information Needed]
Speeds, Sizes, Times [optional]
[More Information Needed]
Evaluation
Testing Data, Factors & Metrics
Testing Data
[More Information Needed]
Factors
[More Information Needed]
Metrics
[More Information Needed]
Results
[More Information Needed]
Summary
Model Examination [optional]
[More Information Needed]
Environmental Impact
Carbon emissions can be estimated using the
Machine Learning Impact calculator presented in
Lacoste et al. (2019).
- Hardware Type: [More Information Needed]
- Hours used: [More Information Needed]
- Cloud Provider: [More Information Needed]
- Compute Region: [More Information Needed]
- Carbon Emitted: [More Information Needed]
Technical Specifications [optional]
Model Architecture and Objective
[More Information Needed]
Compute Infrastructure
[More Information Needed]
Hardware
[More Information Needed]
Software
[More Information Needed]
Citation [optional]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
Glossary [optional]
[More Information Needed]
More Information [optional]
[More Information Needed]
Model Card Authors [optional]
[More Information Needed]
Model Card Contact
[More Information Needed]