Frozen v2 chip could dramatically reduce operating costs by generating six to ten times more AI tokens using the same electricity as current TPUs
Too little corroboration in the last 3 days to call a trend (3 articles). Watching for it to gain traction.
"The chip would reportedly embed portions of the Gemini architecture directly into silicon... At six times the efficiency, the power required for a given volume of tokens could theoretically fall by about 83%. At ten times, the reduction could reach 90%."
"Google engineers believe Frozen v2 could generate six to ten times more AI tokens using the same amount of electricity compared with Google's latest Tensor Processing Units (TPUs). Producing more tokens with the same power means better efficiency and lower operating costs."
"the chip could be 6 to 10 times more efficient than Google's latest custom Tensor Processing Units (TPUs), measured by AI tokens served per unit of power"