Robots in the real world encounter new tools and constraints daily. Karla allows an agent to learn a new physical tool or API on the fly during inference, modifying its own weights in seconds on consumer hardware (e.g., RTX 4060 Ti), without ever forgetting its foundational knowledge.
Instead of a standard transformer pipeline, Karla acts as a multi-frequency brain based on the Nested Learning (NL) paradigm [1]:
1pip install torch transformers datasets pandas accelerate bitsandbytes
2cd Karla
3python chat.py
4
[1] Behrouz, A., Razaviyayn, M., Zhong, P., & Mirrokni, V. (2025).
Nested Learning: The Illusion of Deep Learning Architecture.
Google Research.
arXiv:2512.24695
[2] Darlow, L., Regan, C., Risi, S., Seely, J., & Jones, L. (2025).
Continuous Thought Machines.
Sakana AI.
arXiv:2505.05522