This repository contains the cutting-edge Aura-1-Thinking reasoning model available as a standard base layer for native PyTorch environments.
This engine is tailor-made to deliver high-speed token streaming, running smoothly on accessible hardware layouts including standard cloud instances with 16GB RAM.
Performance Footprint
Baseline Raw (FP16):~8.5 GB - 10.0 GB | Standard CPU / Basic GPU Space