Views
No views yet
Run an 8-billion-parameter 1-bit LLM in 1.1 GB on a $99 Jetson Nano.
| Model | Size on disk | RAM used | Prompt | Generation | Board |
|---|---|---|---|---|---|
| Bonsai-8B | 1.1 GB | 2.5 GB | 2.1 tok/s | 1.1 tok/s | Jetson Nano 4GB |
| Bonsai-4B | 546 MB | ~1.5 GB | 3.6 tok/s | 1.6 tok/s | Jetson Nano 4GB |
if constexpr, std::is_same_v, structured bindings, fold expressionsnv_bfloat16 type stub, cooperative_groups/reduce.h, CUDA_R_16BFvld1q_*_x* intrinsics-lstdc++fs for std::filesystembinbcast.cu fold expression silently computing nothingCUDA_STANDARD 14, flash attention template exclusionbinbcast.cu was replaced with (void)0. This silently broke ALL binary operations (add, multiply, subtract, divide). The model loaded, allocated memory, ran inference — and produced complete garbage. The fix was one line.