POCKET-35B claims 59 tok/s on CPU and runs locally on phones without a GPU
Powerful_Evening5495 · reddit · 2026-07-26
POCKET-35B is being pitched as a 35B model that runs locally on PCs and even phones with no GPU, using stock llama.cpp and no CUDA or cloud setup.
- The lineup includes 21 GB Q4, 13 GB Q2, 8.2 GB IQ1, and mobile-focused variants
- Claimed targets range from 32 GB RAM PCs to Android phones and Apple devices
- The post highlights CPU-only throughput of 59 tokens/s and a Korean-tuned variant
More from Infra
- They analyzed context from over 1 million videos for $140 in a day — eptwts · 2026-07-26
- ASML’s EUV tools keep Europe in the AI race beyond the model layer — kimmonismus · 2026-07-26
- Combined MCP server adds Redshift querying and S3 Markdown semantic search — modelcontextprotocol · 2026-07-26
- GitHub paid $100,000 for a critical bug in its internal Git infrastructure — sharpeye_wnl · 2026-07-26
- Longsys reaches 7.67% DRAM share as China’s fourth memory player, but HBM remains the hard wall — 量子位 · 2026-07-26
- Is a 64GB M4 Max MacBook Pro the best value for running local models? — guyastronomer · 2026-07-26