logo
  • menu
  • Markets
  • ETFs
  • Live
  • Spot
  • Futures
  • Bots
  • Learn
  • Sign In
  • Sign Up
  • Downloads
  • English
  • |
  • USD
  • |
Sign Up
Crypto PricesLearnLatest NewsDownloadsMarketsSpotAnnouncements
Home/
Latest News/
Live

Google Found a Way to Make Local AI Up to 3x Faster—No New Hardware Required

By Decrypt
May 7, 2026
4.5 
★
★
★
★
★
★
★
★
★
★
 160 User Rating
Share

Running an AI model on your own computer is great—until it isn't.

The promise is privacy, no subscription fees, and no data leaving your machine. The reality, for most people, is watching a cursor blink for five seconds between sentences.

That bottleneck has a name: inference speed. And it has nothing to do with how smart the model is. It's a hardware problem. Standard AI models generate text one word fragment—called a token—at a time. The hardware has to shuttle billions of parameters from memory to its compute units just to produce each single token. It's slow by design. On consumer hardware, it's painful.

The approach is called speculative decoding, and it's been around as a concept for years. Google researchers published the foundational paper back in 2022. The idea didn't go mainstream until now because it required the right architecture to make it work at scale.

Here's the short version of how it works. Instead of making the big, powerful model do all the work alone, you pair it with a tiny "drafter" model. The drafter is fast and cheap—it predicts several tokens at once in less time than the main model would take to produce just one. Then the big model checks all of those guesses in a single pass. If the guesses are right, then you get the whole sequence for the price of one forward pass.

Nothing is sacrificed: The large model—Gemma 4's 31B dense version, for example—still verifies every token, and the output quality is identical. You're just exploiting idle compute power that was sitting unused during the slow parts.

Google says the drafter models share the target model's KV cache—a memory structure that stores already-processed context—so they don't waste time recalculating things the larger model already knows. For the smaller edge models designed for phones and Raspberry Pi devices, the team even built an efficient clustering technique to further cut generation time.

This isn't the only attempt the AI world has made at parallelizing text generation. Diffusion-based language models—like Mercury from Inception Labs—tried a completely different approach: Instead of predicting one token at a time, they start with noise and iteratively refine the entire output. That’s fast on paper, but diffusion LLMs have struggled to match the quality of traditional transformer models, leaving them more of a research curiosity than a practical tool.

Speculative decoding is different because it doesn't change the underlying model at all. It's a serving optimization, not an architecture replacement. The same Gemma 4 you'd already run gets faster.

The practical upside is real. A Gemma 4 26B model running on an Nvidia RTX Pro 6000 desktop GPU gets roughly twice the tokens per second with the MTP drafter enabled, according to Google's own benchmarks. On Apple Silicon, batch sizes of 4 to 8 requests unlock around 2.2x speedups. Not quite the 3x ceiling in every scenario, but still a meaningful difference between "barely usable" and "actually fast enough to work with."

Chrome Is Quietly Installing a 4GB AI Model on Your Computer—And Putting It Back If You Delete It

Google says the drafter unlocks "improved responsiveness: drastically reduce latency for near real-time chat, immersive voice applications and agentic workflows"—the kind of tasks that demand low latency to feel useful at all.

Use cases snap into focus quickly: A local coding assistant that doesn't lag; a voice interface that responds before you've forgotten what you asked; an agentic workflow that doesn't make you wait three seconds between steps. All of this, on hardware you already own.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of BitKan. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. BitKan shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. Products mentioned in this article may not be available in your region.

Latest News

Industry

Cryptocurrency

Airdrop

Markets

  • Brazil’s CVM Launches 60-Day Sprint to Tokenize Securities

    Brazil’s CVM Launches 60-Day Sprint to Tokenize Securities

    The Brazilian Securities and Exchange Commission (CVM) has officially established a dedicated task force to develop an experimental regulatory framework for tokenized securities, providing a fast-tracked timeline for the digital capital market.
    Martha Grizzard
    Jul 21, 2026
  • Hyperliquid Enables Permissionless Markets With HIP-4 Plan

    Hyperliquid Enables Permissionless Markets With HIP-4 Plan

    Hyperliquid has announced a forthcoming enhancement to its HIP-4 upgrade that will allow for the permissionless deployment of decentralized prediction markets.
    Christopher Smith
    Jul 21, 2026
  • DTCC Launches Live Tokenized Asset Trading for Wall Street

    DTCC Launches Live Tokenized Asset Trading for Wall Street

    The DTCC successfully transitioned from pilot testing to live production trades on July 15, 2026, marking the largest-scale institutional tokenization initiative to date.
    Cornell Rachel
    Jul 16, 2026
  • South Korea Updates Asset Law to Include Cryptocurrency

    South Korea Updates Asset Law to Include Cryptocurrency

    The South Korean Ministry of Economy and Finance announced a transition from the 1950 State Property Act to a new National Asset Basic Act to better reflect modern digital resources.
    Martha Grizzard
    Jul 16, 2026
  • New SEC Crypto Rule to Cut Red Tape for Startup Fundraising

    New SEC Crypto Rule to Cut Red Tape for Startup Fundraising

    The U.S. Securities and Exchange Commission plans to introduce a major regulatory framework this month to simplify capital formation and reduce operational hurdles for cryptocurrency businesses.
    Martha Grizzard
    Jul 8, 2026
View more data 
BTCBTC(BTC)
$0
--(Last 24h)
SpotFutures

Top

View more
  1. 1S&P 500 Reclaims 200-Day Moving Average, Bitcoin Gains
  2. 2Trump Softens His Stance on Reciprocal Tariffs, US Stocks and Crypto Markets Rise
  3. 3Vitalik Buterin : The current price of ETH has not been affected by the merger event
  4. 4Vibhu Norby : Solana Spaces store to bring 100K people to Solana per month
  5. 5CZ: compared with the record high nine months ago, the current situation of the industry is much better

Top Gainers

View more
Grvt
GrvtGRVT

$0.2688

+437.66%
Koma Inu
Koma InuKOMA

$0.0240

+84.95%
Kekius Maximus
Kekius MaximusKEKIUS

$0.005282

+61.63%
Unipeg
UnipegUPEG

$568.920

+56.45%
AXT
AXTAXTIB

$58.6000

+43.94%

Top Trending

View more
Strategy
StrategyMSTR

$92.7200

-4.38%
Sei Network
Sei NetworkSEI

$0.0417

-0.57%
Momentum
MomentumMMT

$0.2406

+15.45%
Giggle Fund
Giggle FundGIGGLE

$36.9300

+30.82%
Enso
EnsoENSO

$0.8980

+5.77%

Recently added

View more
Grvt
GrvtGRVT

$0.2687

+437.46%
Direxion Semiconductor Bear 3X ETF
Direxion Semiconductor Bear 3X ETFSOXSB

$52.3300

-16.75%
VanEck Semiconductor ETF
VanEck Semiconductor ETFSMHB

$547.600

+4.01%
PayPal
PayPalPYPLB

$57.4800

-0.43%
Goldman Sachs
Goldman SachsGSB

$1,035.01

+3.81%

Learn

View more
  1. 1What Are ARC-20 Tokens? How Do ARC-20 Tokens Work?
  2. 2What Is the Usual Protocol? How Does Its Tokenomics Work?
  3. 3What Are AI Agent Frameworks? How Do They Power Cryptocurrency?
  4. 4What Is JPYSC? How Japan’s Regulated Stablecoin Works
  5. 5Are AI Agents Safe for Crypto? How to Secure Your Assets
About Us
  • About BitKan
  • Contact Us
  • Announcements
  • VIP Program
  • BitKan Ambassador
  • Institutional Services
Products
  • Spot
  • Futures
  • Crypto Prices
  • Learn
  • News
  • Markets
  • How to Buy Crypto
  • BTC to USD Calculator
  • Reward
Help
  • Help Center
  • Email Us
  • Live Chat
  • Download APP
  • Listing Application
  • Buy Bitcoin
  • Buy Ethereum
  • Buy Dogecoin
  • Buy Altcoins
Terms
  • Terms of Use
  • Privacy Policy
  • Trading Rules
  • Fee
K-Site
English
About Us
+
  • About BitKan
  • Contact Us
  • Announcements
  • VIP Program
  • BitKan Ambassador
  • Institutional Services
Products
+
  • Spot
  • Futures
  • Crypto Prices
  • Learn
  • News
  • Markets
  • How to Buy Crypto
  • BTC to USD Calculator
  • Reward
Help
+
  • Help Center
  • Email Us
  • Live Chat
  • Download APP
  • Listing Application
  • Buy Bitcoin
  • Buy Ethereum
  • Buy Dogecoin
  • Buy Altcoins
Terms
+
  • Terms of Use
  • Privacy Policy
  • Trading Rules
  • Fee
K-Site
+
  • Twitter
  • Facebook
  • Telegram
  • YouTube
  • Instagram
  • Medium
  • Linkedin
@2012-2026 BITKAN.com