logo
  • menu
  • Markets
  • ETFs
  • Live
  • Spot
  • Futures
  • Bots
  • Learn
  • Sign In
  • Sign Up
  • Downloads
  • English
  • |
  • USD
  • |
Sign Up
Crypto PricesLearnLatest NewsDownloadsMarketsSpotAnnouncements
Home/
Latest News/
Live

Meet Bonsai: The First 27B AI Model That Fits on Your Phone

By Decrypt
Jul 16, 2026
4.5 
★
★
★
★
★
★
★
★
★
★
 249 User Rating
Share

I models eat up a lot of memory. A 27-billion-parameter AI model, considered medium-sized by industry standards, needs roughly 54 GB of memory to run on half precision. Most laptops can't hold that. Some desktop rigs can't either.

Earlier this week, PrismML released one at 3.9 GB—small enough to fit on an iPhone.

Parameters are the number of dials and tweaks a model can handle. The more parameters, the denser and more capable a model is.

The compression method, built on Caltech intellectual property, reduces each model weight from 16 bits of floating-point precision to a single sign—+1 or -1 in the binary build, one of three values in the ternary. Each group of 128 weights shares a 16-bit scaling factor, landing the binary variant at 1.125 bits per weight: 14 times smaller than the full-precision original. The ternary model adds a zero state for slightly more expressive power and settles at 1.71 bits.

In easier terms, this means a ternary AI model uses only three settings for each internal value—negative, zero, or positive—while a standard AI can choose from about 65,000 settings.

PrismML did that without losing much of the output quality.

What makes this different from conventional "low-bit" models is that nothing gets a higher-precision escape hatch: embeddings, attention, and the full language model head are all compressed end-to-end. Most quantized builds keep certain sensitive layers at full precision, which ends up increasing their size as a tradeoff for better quality. Bonsai doesn't play that game.

Benchmarks

Across 15 benchmarks evaluated in thinking mode on NVIDIA H100 GPUs—spanning knowledge, math, coding, and tool use—Ternary Bonsai 27B averages 80.49, or 94.6% of the full-precision model. The 1-bit variant hits 76.11.

Overall, on benchmarks, the models perform much better than Gemma 4 or Qwen 3.6 in terms of how much potential they offer for their size.

The models are pretty good for what they offer, and considering how little resources they require, they take small hardware (smartphones and lower end PCs) to another level in terms of capabilities. AIME25 and AIME26, modeled on the American Invitational Mathematics Examination, come in 93.7% for Ternary Bonsai 27B versus 95.3% for the much bigger Qwen 3.6B. Bonsai scores 86 points in codig vs 88 for Qwen 3.6 and 77% on general knowledge vs 83 for Qwen 3.6.

The model also uses a hybrid attention backbone where roughly 75% of the layers are linear rather than full quadratic attention. That architecture is what makes a 262K-token context window practical on-device—something a standard attention stack would make prohibitively expensive on phone hardware.

We tested it

We ran Bonsai 27B ourselves. Coding takes iteration: single-shot prompts won't compete with cloud frontier models. Being local and free makes that irrelevant. For our Zombie Type game—a first-person typing-horror browser game—two vibe coding rounds produced clean collision detection, proper scoring logic, and graphics that held together. The model grasps structure early; the second pass refines rather than rebuilds.

Interestingly enough, some models (like the skeletons) looked more elaborate than the ones from GPT 5.6 Sol. It doesn’t mean it’s better by any means, just that on this task it produced a cute skeleton whereas the AI king made a poorer stylistic choice.

Creative writing is a more qualified story, and the criteria is more subjective.

Roughly speaking, the results aren't particularly imaginative if you have a zero-shot prompt in mind.

That said, Bonsai produces stories with consistent internal logic, pacing, and arc—better, or on par with Claude Haiku or even Sonnet on lower effort on comparable prompts. For a model that runs entirely on your own hardware with no API costs, that's a lot to say.

PrismML also ships a DSpark speculative decoding layer alongside the model—a lightweight drafter that proposes blocks of candidate tokens, which the main model verifies in a single forward pass rather than generating token-by-token. On an H100 that adds a 1.37x throughput boost with no change in output quality, since verification preserves the exact output distribution. On Apple Silicon it's not yet enabled by default, but for GPU serving it's a real gain.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of BitKan. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. BitKan shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. Products mentioned in this article may not be available in your region.

Latest News

Industry

Cryptocurrency

Airdrop

Markets

  • Brazil’s CVM Launches 60-Day Sprint to Tokenize Securities

    Brazil’s CVM Launches 60-Day Sprint to Tokenize Securities

    The Brazilian Securities and Exchange Commission (CVM) has officially established a dedicated task force to develop an experimental regulatory framework for tokenized securities, providing a fast-tracked timeline for the digital capital market.
    Martha Grizzard
    Jul 21, 2026
  • Hyperliquid Enables Permissionless Markets With HIP-4 Plan

    Hyperliquid Enables Permissionless Markets With HIP-4 Plan

    Hyperliquid has announced a forthcoming enhancement to its HIP-4 upgrade that will allow for the permissionless deployment of decentralized prediction markets.
    Christopher Smith
    Jul 21, 2026
  • DTCC Launches Live Tokenized Asset Trading for Wall Street

    DTCC Launches Live Tokenized Asset Trading for Wall Street

    The DTCC successfully transitioned from pilot testing to live production trades on July 15, 2026, marking the largest-scale institutional tokenization initiative to date.
    Cornell Rachel
    Jul 16, 2026
  • South Korea Updates Asset Law to Include Cryptocurrency

    South Korea Updates Asset Law to Include Cryptocurrency

    The South Korean Ministry of Economy and Finance announced a transition from the 1950 State Property Act to a new National Asset Basic Act to better reflect modern digital resources.
    Martha Grizzard
    Jul 16, 2026
  • New SEC Crypto Rule to Cut Red Tape for Startup Fundraising

    New SEC Crypto Rule to Cut Red Tape for Startup Fundraising

    The U.S. Securities and Exchange Commission plans to introduce a major regulatory framework this month to simplify capital formation and reduce operational hurdles for cryptocurrency businesses.
    Martha Grizzard
    Jul 8, 2026
View more data 
BTCBTC(BTC)
$0
--(Last 24h)
SpotFutures

Top

View more
  1. 1S&P 500 Reclaims 200-Day Moving Average, Bitcoin Gains
  2. 2Trump Softens His Stance on Reciprocal Tariffs, US Stocks and Crypto Markets Rise
  3. 3Vitalik Buterin : The current price of ETH has not been affected by the merger event
  4. 4Vibhu Norby : Solana Spaces store to bring 100K people to Solana per month
  5. 5CZ: compared with the record high nine months ago, the current situation of the industry is much better

Top Gainers

View more
RSK Infrastructure Framework
RSK Infrastructure FrameworkRIF

$0.1282

+81.84%
Akedo
AkedoAKE

$0.002297

+34.18%
Elastos
ElastosELA

$0.3200

+27.29%
Prom
PromPROM

$1.8230

+24.61%
Lorenzo Protocol
Lorenzo ProtocolBANK

$0.2936

+22.33%

Top Trending

View more
DeXe
DeXeDEXE

$1.6640

-61.16%
Litecoin
LitecoinLTC

$47.1400

-0.30%
Dogecoin
DogecoinDOGE

$0.0692

-5.17%
Sui Network
Sui NetworkSUI

$0.7419

-3.01%
Ethereum
EthereumETH

$1,870.73

-3.31%

Recently added

View more
Ket
KetKET

$0.0113

+20.36%
Direxion MU Bull 2X ETF
Direxion MU Bull 2X ETFMUUB

$36.1000

-0.52%
GraniteShares 2X Long INTC ETF
GraniteShares 2X Long INTC ETFINTWB

$26.9100

-1.43%
AXT
AXTAXTIB

$53.0000

-3.76%
GraniteShares 2X Long MRVL ETF
GraniteShares 2X Long MRVL ETFMVLLB

$26.1900

-8.62%

Learn

View more
  1. 1What Are AI Agent Frameworks? How Do They Power Cryptocurrency?
  2. 2What Is JPYSC? How Japan’s Regulated Stablecoin Works
  3. 3Are AI Agents Safe for Crypto? How to Secure Your Assets
  4. 4What Is Cross-Chain Interoperability? How Does It Function?
  5. 5What is OUSD? How Does Open USD Work for Digital Payments?
About Us
  • About BitKan
  • Contact Us
  • Announcements
  • VIP Program
  • BitKan Ambassador
  • Institutional Services
Products
  • Spot
  • Futures
  • Crypto Prices
  • Learn
  • News
  • Markets
  • How to Buy Crypto
  • BTC to USD Calculator
  • Reward
Help
  • Help Center
  • Email Us
  • Live Chat
  • Download APP
  • Listing Application
  • Buy Bitcoin
  • Buy Ethereum
  • Buy Dogecoin
  • Buy Altcoins
Terms
  • Terms of Use
  • Privacy Policy
  • Trading Rules
  • Fee
K-Site
English
About Us
+
  • About BitKan
  • Contact Us
  • Announcements
  • VIP Program
  • BitKan Ambassador
  • Institutional Services
Products
+
  • Spot
  • Futures
  • Crypto Prices
  • Learn
  • News
  • Markets
  • How to Buy Crypto
  • BTC to USD Calculator
  • Reward
Help
+
  • Help Center
  • Email Us
  • Live Chat
  • Download APP
  • Listing Application
  • Buy Bitcoin
  • Buy Ethereum
  • Buy Dogecoin
  • Buy Altcoins
Terms
+
  • Terms of Use
  • Privacy Policy
  • Trading Rules
  • Fee
K-Site
+
  • Twitter
  • Facebook
  • Telegram
  • YouTube
  • Instagram
  • Medium
  • Linkedin
@2012-2026 BITKAN.com