Google's TurboQuant compresses the KV cache of large language models to 3 bits. Tests on an AMD card show: The core promise is true — with side effects.
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.
Google's TurboQuant compresses the KV cache of large language models to 3 bits. Tests on an AMD card show: The core promise is true — with side effects.