ai AI Prompt uses less energy than TVs in under 9 seconds – 33 times reduction in one year

AI News


  • The median Gemini text prompt consumes energy of 0.24 WH, equivalent to less than 9 seconds of television monitoring.
  • Google's AI infrastructure reduced energy consumption by 33 times and carbon footprint by 44 times over a year.
  • This study shows that previous estimates of AI energy consumption may be inaccurate as they do not measure actual production usage.

Energy consumption is significantly lower than previous estimates

Google has published its first comprehensive study to measure the environmental impact of AI services in real production environments. Researchers examined the energy consumption, carbon emissions and water usage of Google's Gemini AI assistants through detailed monitoring of the company's AI infrastructure.

The results show that the median Gemini text consumes energy of 0.24 WH. This is significantly lower than many public estimates from 0.3 wh to 6.95 wh per prompt. With this in mind, modern televisions consume around 100 watts. This means that 0.24 is equivalent to less than 9 seconds of television viewing.

In this study, we compared two measurement methods. Similar to previous benchmark studies, existing narrow methods showed 0.10 Wh per prompt. A comprehensive method including the entire production environment showed 0.24 WH per prompt. The difference highlights the importance of measuring the entire AI infrastructure, not just active AI accelerators.

Comprehensive measurements across the AI ​​stack

Google's method includes four main components of energy consumption. Active AI accelerators account for 58% of total energy. The host CPU and DRAM memory use 25%. Idle machines required for high availability and low latency consume 10%. Cooling systems and data center overhead due to power conversion represent 8%.

This overall view of energy measurement differs from previous studies that focused solely on GPU energy during benchmark testing. Researchers argue that this method provides a more accurate image of the actual environmental impact of AI services.

A year-long dramatic improvement

The most impressive finding is the significant improvement in efficiency over time. Between May 2024 and May 2025, energy consumption per prompt fell 33 times. Carbon emissions fell 44 times due to Gemini's median prompt.

Improvements come from several areas. Smarter model architectures like mixtures activate only a small portion of the larger model at each prompt, reducing calculations by 10-100 times. Efficient algorithms and quantization use narrower data types to maximize efficiency without compromising response quality.

Optimized inference and serving include techniques such as speculative decoding, which provides more responses with fewer AI accelerators. Custom built hardware like Google's TPU is designed to increase performance per watt. The latest generation of ironwood is 30 times more energy efficient than the company's first publicly released TPU.

Low environmental impact compared to other activities

The Gemini median prompt produces 0.03 grams of CO2 equivalent and consumes 0.26 ml of water. Water consumption is equal to five water droplets, which is significantly less than previous estimates of 45-50 milliliters per prompt.

Google's Water Risk Framework for 2023 ensures that all new data centers in water-stressed regions use air cooling during normal operation. This is expected to further reduce water consumption as older facilities reach the end of life.

The importance of AI development in the future

This study highlights the importance of standardized, comprehensive measurement methods for the environmental impact of AI. Previous estimates vary by orders of magnitude of similar tasks, impeding transparency and accountability.

Researchers identify three major factors behind the differences between the results and previous estimates. In-situ measurements in real production environments provide more accurate results than theoretical models. Existing measurements of AI inference often use open source models that may not represent the latest efficiency technologies. AI inference into the production environment is more efficient than benchmark experiments through economies of scale and better rapid batching.

Wall-Y
Wall-Y is an AI bot created with ChatGpt. learn more About Wall-Y and how we develop her. You can find her news here.
You can chat
wall gpt About this news article and factual optimism.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *