OmniNeural — World’s First NPU-aware Multimodal Model
Overview
OmniNeural is the first fully multimodal model designed specifically for Neural Processing Units (NPUs). It natively understands text, images, and audio, and runs across PCs, mobile devices, automobile, IoT, and robotics.
Demos
📱 Mobile Phone NPU - Demo on Samsung S25 Ultra
The first-ever fully local, multimodal, and conversational AI assistant that hears you and sees what you see, running natively on Snapdragon NPU for long battery life and low latency.
✨ PC NPU - Capabilities Highlights
🖼️ Multi-Image Reasoning Spot the difference across two images in multi-round dialogue.
🤖 Image + Text → Function Call Snap a poster, add a text instruction, and AI agent creates a calendar event.
🎶 Multi-Audio Comparison Tell the difference between two music clips locally.
Key Features
Multimodal Intelligence – Processes text, image, and audio in a unified model for richer reasoning and perception.
NPU-Optimized Architecture – Uses ReLU ops, sparse tensors, convolutional layers, and static graph execution for maximum throughput — 20% faster than non-NPU-aware models .
Hardware-Aware Attention – Attention patterns tuned for NPU, lowering compute and memory demand .
This model is released under the Creative Commons Attribution–NonCommercial 4.0 (CC BY-NC 4.0) license.
Non-commercial use, modification, and redistribution are permitted with attribution.
For commercial licensing, please contact dev@nexa.ai.