AIToday

Qwen 35B model runs on mobile device at 90 tokens/sec input speed

r/MachineLearning1d ago

Key takeaway

A researcher demonstrated that a 35-billion-parameter language model can run on a mobile S26 Ultra device while maintaining precision, achieving 90 tokens per second for input processing and 8 tokens per second for output. The work is part of the researcher's self-directed AI learning and represents a potential shift toward running large models directly on consumer phones rather than relying solely on cloud servers.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A researcher tested Qwen's 35B mixture-of-experts language model on an S26 Ultra mobile device and found the active model footprint fits within the device's memory limits, with early testing showing roughly 90 input processing tokens per second after optimization and output generation around 8 tokens per second.

  • Why it matters

    Running a 35-billion-parameter AI model on consumer mobile hardware without degrading precision suggests that capable language models may become deployable on phones rather than requiring cloud infrastructure. The researcher learned AI and machine learning independently (without a PhD) and has access to computing resources to validate this approach.

  • What to watch

    The researcher is seeking collaborators to expand testing across devices. Four papers authored by this researcher remain under review on arXiv; institutional affiliation may affect their publication prospects.

In Depth

The researcher began testing Qwen's 35B mixture-of-experts language model—a large language model with 35 billion parameters—on an S26 Ultra mobile device. The core finding is that the active model footprint (the memory required to run the model in practice) fits within the device's available memory. After optimization, the model achieves roughly 90 tokens per second for input processing—the speed at which the model reads and understands incoming text—and approximately 8 tokens per second for output generation, the speed at which it produces responses. No precision loss is reported, meaning the model's accuracy was not sacrificed to achieve these speeds. The researcher notes that the methods and architecture used are not being shared publicly at this stage. The post also includes personal context: the researcher came to AI and machine learning through self-directed study and personal interest rather than formal graduate training, and has invested in computing resources to conduct this kind of testing. The researcher mentions having submitted four papers to arXiv (a preprint repository widely used in machine learning) as first author, but all remain under review; the researcher attributes the hold-up to their lack of institutional affiliation, a barrier that can affect publication timelines in academic contexts. The post closes with an invitation for collaborators interested in expanding the testing effort.

Context & Analysis

This post represents a case of independent AI research operating outside traditional institutional frameworks. The researcher explicitly notes having learned machine learning based on personal interest rather than through formal PhD study, and has accessed sufficient compute and resources to test a large language model on mobile hardware. The constraint mentioned—that four papers remain on hold on arXiv due to first-authorship status and lack of institutional affiliation—reflects a structural tension in academic publishing: peer review systems often advantage researchers with formal institutional backing, even when the technical work itself may be sound. The call for collaborators suggests the researcher views this as a proof-of-concept worth validating more broadly, though the decision not to share methods or architecture details limits immediate reproducibility.

FAQ

What model and device were used in the test?
The test used Qwen's 35B mixture-of-experts language model on an S26 Ultra mobile device.
What are the performance speeds achieved?
Early testing shows roughly 90 input processing tokens per second after optimization and output generation around 8 tokens per second on the mobile device.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →