This technical report describes Hayula Labs' fleet strategy for orchestrating AI inference across a heterogeneous collection of consumer-grade machines. We detail the architecture, operational metrics, and cost analysis of a production deployment spanning five machines in Kuwait (Mac Studio M2 Ultra 192GB, Acer Nitro V14 with RTX 4050 6GB, Framework Desktop, legacy workstations, NAS) with a unified routing layer. Over 90 days of operation, the fleet processed 3.9M inference requests with 99.4% u
1@techreport{hayulalab2026fleetstrategy,
2 title={Technical Report: The Fleet Strategy — Distributed Inference Across Heterogeneous Consumer Hardware — Hayula Research},
3 author={Hayula AI Lab},
4 year={2026},
5 url={https://huggingface.co/hayulalab/fleet-strategy-paper}
6}