AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Developers have released Maple-Preview, a Ternary 20-billion-parameter Mixture of Experts model, capable of processing 120 tokens per second on an iPhone. This highlights advancements in mobile AI performance.

Developers have unveiled Maple-Preview, a Ternary 20-billion-parameter Mixture of Experts (MoE) model capable of running at 120 tokens per second on an iPhone. This marks a significant milestone in mobile AI performance, demonstrating that large-scale models can operate efficiently on consumer smartphones.

The project, shared via Show HN, showcases Maple-Preview’s ability to process natural language inputs rapidly on an Apple iPhone, leveraging a Ternary MoE architecture. The model’s inference speed of 120 tokens/sec was confirmed by the developer, indicating promising scalability for AI applications on mobile devices.

According to the developer, Maple-Preview is based on a Ternary 20B MoE, which uses three possible states per parameter, reducing computational complexity compared to traditional dense models. The demonstration was achieved without specialized hardware, relying solely on an iPhone, suggesting potential for widespread adoption.

At a glance
announcementWhen: announced March 2024
The developmentMaple-Preview, a new AI model, demonstrates high-speed inference of a Ternary 20B MoE on an iPhone, emphasizing progress in mobile AI deployment.

Implications for Mobile AI Deployment

This development signifies a breakthrough in deploying large language models on consumer-grade smartphones. Achieving 120 tokens/sec inference on an iPhone indicates that advanced AI capabilities could become more accessible, enabling new applications in real-time language processing, virtual assistants, and on-device AI services without relying on cloud computing.

It also challenges previous assumptions that such large models require extensive hardware or cloud infrastructure, opening pathways for more privacy-preserving, low-latency AI solutions directly on devices.

Amazon

iPhone compatible AI language model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Advances in Ternary MoE and Mobile AI

Recent years have seen rapid progress in Mixture of Experts (MoE) architectures, which enable large models with fewer parameters to operate efficiently by activating only parts of the network as needed. Ternary models, which use three states per parameter, further reduce computational demands.

Prior to this, large language models like GPT-3 and GPT-4 have primarily run on specialized hardware or cloud servers, limiting on-device use. The demonstration of Maple-Preview on an iPhone marks a shift toward more capable mobile AI, driven by innovations in model architecture and optimization techniques.

“This is the first time we’ve seen a Ternary 20B MoE run at such speed on a standard iPhone, opening new possibilities for on-device AI.”

— Developer of Maple-Preview

Amazon

mobile AI processing device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Model Performance and Scalability

While inference speed has been demonstrated, details about the model’s accuracy, robustness, and energy consumption on the iPhone remain unclear. It is also not yet confirmed how well the model performs across diverse tasks or in real-world scenarios.

Further testing is needed to verify whether this performance can be sustained in continuous use or with larger input sizes, and whether similar results can be replicated across different mobile devices.

Amazon

on-device AI assistant app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Mobile AI Model Development

Developers and researchers are expected to explore optimizing Maple-Preview further, testing its capabilities across various applications, and assessing its performance in real-world settings. Additionally, efforts will likely focus on improving model accuracy, reducing energy consumption, and expanding compatibility to other mobile platforms.

Further releases or updates may showcase larger models or enhanced versions, pushing the boundaries of what is possible with AI on smartphones.

Amazon

AI model for iPhone

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Maple-Preview achieve such high speed on an iPhone?

It uses a Ternary 20B Mixture of Experts architecture, which reduces computational complexity by employing three states per parameter, allowing faster inference without specialized hardware.

Can this model run other AI tasks besides language processing?

While primarily demonstrated for language, the architecture’s efficiency suggests potential for other tasks, but specific applications have not yet been confirmed.

Does this mean large AI models will soon be available on all smartphones?

This development indicates progress toward that goal, but widespread deployment depends on further optimization, robustness, and energy efficiency improvements.

What are the limitations of Maple-Preview currently?

Details about the model’s accuracy, long-term stability, and energy usage on mobile devices are still unclear, and further testing is needed to confirm its practical viability.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: Lathe – Use LLMs to learn a new domain, not skip past it

Lathe is a new tool that generates interactive, multi-part tutorials for learning technical skills, aiming to enhance hands-on learning with LLMs.

The Local-First Agentic Operator

A single operator, empowered by agentic AI, now builds and manages multiple complex products across domains, traditionally requiring organizations.

Creating a Structural Model for a Post-Ai Civilization

Many envision a post-AI civilization’s future, but crafting a resilient, equitable structure requires careful planning and innovative foresight.

The 2028 Model Lab Endgame: How Six Becomes Two, Three, or Twelve

A 2026 forecast predicts three possible futures for Western frontier AI labs by 2028, with significant implications for investment and strategy.