TL;DR
Developers have released Maple-Preview, a Ternary 20-billion-parameter Mixture of Experts model, capable of processing 120 tokens per second on an iPhone. This highlights advancements in mobile AI performance.
Developers have unveiled Maple-Preview, a Ternary 20-billion-parameter Mixture of Experts (MoE) model capable of running at 120 tokens per second on an iPhone. This marks a significant milestone in mobile AI performance, demonstrating that large-scale models can operate efficiently on consumer smartphones.
The project, shared via Show HN, showcases Maple-Preview’s ability to process natural language inputs rapidly on an Apple iPhone, leveraging a Ternary MoE architecture. The model’s inference speed of 120 tokens/sec was confirmed by the developer, indicating promising scalability for AI applications on mobile devices.
According to the developer, Maple-Preview is based on a Ternary 20B MoE, which uses three possible states per parameter, reducing computational complexity compared to traditional dense models. The demonstration was achieved without specialized hardware, relying solely on an iPhone, suggesting potential for widespread adoption.
Implications for Mobile AI Deployment
This development signifies a breakthrough in deploying large language models on consumer-grade smartphones. Achieving 120 tokens/sec inference on an iPhone indicates that advanced AI capabilities could become more accessible, enabling new applications in real-time language processing, virtual assistants, and on-device AI services without relying on cloud computing.
It also challenges previous assumptions that such large models require extensive hardware or cloud infrastructure, opening pathways for more privacy-preserving, low-latency AI solutions directly on devices.

DOSUKE Translation Earbuds, 3-in-1 AI Language Translator Earbuds with Premium Sound, Long Battery Life, Translating Earbud with Charging Case for Business, Learning, and Travel, Modern Black
- Easy Setup: Connect and scan with free AI COOL app
- 3-in-1 Functionality: Translate, listen to music, and make calls
- Active Noise Cancellation: Reduces noise and enhances privacy
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Advances in Ternary MoE and Mobile AI
Recent years have seen rapid progress in Mixture of Experts (MoE) architectures, which enable large models with fewer parameters to operate efficiently by activating only parts of the network as needed. Ternary models, which use three states per parameter, further reduce computational demands.
Prior to this, large language models like GPT-3 and GPT-4 have primarily run on specialized hardware or cloud servers, limiting on-device use. The demonstration of Maple-Preview on an iPhone marks a shift toward more capable mobile AI, driven by innovations in model architecture and optimization techniques.
“This is the first time we’ve seen a Ternary 20B MoE run at such speed on a standard iPhone, opening new possibilities for on-device AI.”
— Developer of Maple-Preview

Innioasis PR1 AI Voice Recorder, Transcription & Translation with AI, 3.99 inch 64GB Smart Summarize, Offline AI Processing Note Taker for Meeting & Lectures, Translation for Business Travel, Black
- Offline Transcription: No internet needed, instant speech-to-text
- Subscription-Free: No monthly fees for transcription services
- Privacy & Security: Voice data remains on device, no cloud uploads
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unconfirmed Aspects of Model Performance and Scalability
While inference speed has been demonstrated, details about the model’s accuracy, robustness, and energy consumption on the iPhone remain unclear. It is also not yet confirmed how well the model performs across diverse tasks or in real-world scenarios.
Further testing is needed to verify whether this performance can be sustained in continuous use or with larger input sizes, and whether similar results can be replicated across different mobile devices.

ZNP Z02 Smart AI Companion Digital Badge, HD Touch Screen Bluetooth 6.0 Translator, No Subscription Wearable AI Assistant with Meeting Minutes, Memo & Custom Wallpaper for Travel Business Daily Use
- All-in-One AI Recorder & Translator: Voice recorder, 102-language translator, meeting assistant
- No Subscription Required: Supports instant translation and high-quality audio recording
- Smart Meeting Assistant: Real-time speaker detection and dual recording modes
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in Mobile AI Model Development
Developers and researchers are expected to explore optimizing Maple-Preview further, testing its capabilities across various applications, and assessing its performance in real-world settings. Additionally, efforts will likely focus on improving model accuracy, reducing energy consumption, and expanding compatibility to other mobile platforms.
Further releases or updates may showcase larger models or enhanced versions, pushing the boundaries of what is possible with AI on smartphones.
As an affiliate, we earn on qualifying purchases.
Key Questions
How does Maple-Preview achieve such high speed on an iPhone?
It uses a Ternary 20B Mixture of Experts architecture, which reduces computational complexity by employing three states per parameter, allowing faster inference without specialized hardware.
Can this model run other AI tasks besides language processing?
While primarily demonstrated for language, the architecture’s efficiency suggests potential for other tasks, but specific applications have not yet been confirmed.
Does this mean large AI models will soon be available on all smartphones?
This development indicates progress toward that goal, but widespread deployment depends on further optimization, robustness, and energy efficiency improvements.
What are the limitations of Maple-Preview currently?
Details about the model’s accuracy, long-term stability, and energy usage on mobile devices are still unclear, and further testing is needed to confirm its practical viability.
Source: hn