PrismML hopes its tiny LLM could change how we all use AI
First reported by TechCrunch ·
AI models can now fit on your phone and run locally, making them faster, cheaper, and more private.
PrismML, a startup founded by Caltech researchers, has released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model. The new model is 5.9 GB, a 9x to 10x reduction, making it small enough to potentially run on PCs and high-end smartphones. This compression technique, utilizing "ternary" weights (values of +1, -1, or 0), allows the model to retain 98% of Qwen's benchmark performance. The company, led by Caltech professor Babak Hassibi and advised by Databricks co-founder Ion Stoica, aims to apply this compression to even larger models, citing potential partnerships with companies like Apple. Previous versions of PrismML's models have seen significant download numbers, indicating market interest in smaller, capable AI.
PrismML's success in drastically reducing LLM size while maintaining high performance signals a significant shift toward edge AI and on-device processing. This approach democratizes access to powerful AI by removing the reliance on cloud infrastructure, potentially lowering costs for both developers and end-users. The ability to run complex models locally also addresses privacy concerns, as data no longer needs to be transmitted to external servers for processing. This could accelerate the adoption of AI in sensitive applications and devices where cloud connectivity is limited or undesirable.
The advancement in LLM compression directly impacts the competitive landscape for AI hardware and software. Companies developing specialized AI chips or optimized inference engines may find their markets challenged by highly efficient software solutions running on general-purpose hardware. This development also sets a new benchmark for performance and efficiency, potentially forcing larger AI players to reconsider their strategies for model deployment and development. Watch for increased investment in on-device AI capabilities and the emergence of new platforms that leverage these smaller, more accessible LLMs.
AI-written summary. May contain errors.