Read that Google Cloud has reportedly been turning down customer deals because it doesn't have enough compute internally, and I immediately wanted to know what they're actually building to fix it. Turns out it's a chip codenamed Frozen V2, and the backstory is genuinely interesting. An earlier version of this idea, led by DeepMind's chief scientist, Jeff Dean, wanted to burn Gemini's actual model weights directly into the silicon. Got scrapped, because the chip would've become useless the moment Gemini updated to a new version. Frozen V2 fixes that by hardwiring only the architecture instead, the stable underlying blueprint, while leaving the weights updatable through normal loading. The reported result is 6 to 10x more tokens processed per watt compared to Google's current TPUs! Targeted for deployment as early as 2028, running alongside the existing TPU lineup rather than replacing it. Feels like the clearest sign yet that the biggest labs are starting to treat the model and the chip as one design problem instead of two separate ones!
NeilXbtShare


Source:Show original
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information.
Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.