Inference Power Draw: 70.0 W (Strict Sub-TDP Cap on NVIDIA T4)
Energy Efficiency: 4.22104 J/tok
Throughput / Latency: ~60.3 ms/tok
Architectural Overview
This adapter integrates a frozen positive-definite metric tensor $G \in [0.1, 1.0]$ across all transformer self-attention query-key projections. By enforcing an invariant geometric manifold during continuous token generation, it bounds non-equilibrium steady-state (NESS) entropy drift, eliminating parameter thrashing and stabilizing long-context energy draw.