Views
No views yet
pivotai-ft (trained on clean synthetic pairs), this model was trained on agent reasoning chains — the hypothesis being that richer teacher signal improves generalization. Results were mixed: reasoning coherence improved, but structural output compliance dropped.| Property | Value |
|---|---|
| Base model | unsloth/Meta-Llama-3.1-8B |
| Training method | QLoRA r=8, α=16, dropout=0.05 |
| Training data | 449 Alpaca-format distillation pairs (Phase 2 agent traces) |
| Epochs | 5 |
| Final train loss | 0.429 |
| Hardware | Lightning.ai A100 (bf16, seq_len=16384) |
| Format | GGUF Q4_K_M (4.6 GB) |
| Metric | Score | Target | ✓/✗ |
|---|---|---|---|
| JSON valid | 92.4% | 85% | ✓ |
| Savings found | 98.1% | 70% | ✓ |
| Budget compliance | — | 80% | — |
| Schema compliance | 0.0% | 80% | ✗ |
| BERTScore F1 | 0.738 | 0.70 | ✓ |
| ROUGE-L | 0.090 | 0.25 | ✗ |
| Reasoning coherence | 0.674 | 0.65 | ✓ |
| Grounding accuracy | 0.442 | 0.60 | ✗ |
| Red-team pass | 46.7% | 80% | ✗ |
1ollama create pivotai-distill -f Modelfile.distill
2ollama run pivotai-distill### Instruction:
Act as pivotai Supervisor for an Indian domestic trip. Coordinate the Analyst, Concierge, and Optimizer agents to find Price-Pivot Points and produce an optimized itinerary. Show the reasoning chain for each agent handoff, then provide the final pivot analysis and optimized itinerary.
### Input:
{"starting_city": "Mumbai", ...}
### Response:Patnaik, A. V. S. (2026). Cost-Matched Data Generation for LLM Fine-Tuning: Comparing
Supervised Fine-Tuning, Knowledge Distillation, and Curriculum Learning for an Agentic
Travel-Planning System. Zenodo. https://doi.org/10.5281/zenodo.21198884