This technical report bridges the gap between textbook large language model (LLM) theory and practical deployment on consumer-grade Apple Silicon hardware. We apply key results from Foundations of Large Language Models (Xiao & Zhu, 2025) to the specific constraints and capabilities of the Mac Studio M2 Ultra (192 GB unified memory, 76 GPU cores) running in the Hayula AI ecosystem. We present actionable guidance across five dimensions: (1) inference optimization — KV cache management, continuous
1@techreport{hayulalab2026llmfoundationsapplied,
2 title={LLM Foundations Applied: Optimizing Large Language Models for Apple M2 Ultra Consumer Hardware — Hayula Research},
3 author={Hayula AI Lab},
4 year={2026},
5 url={https://huggingface.co/hayulalab/llm-foundations-applied-paper}
6}