As we approach the end of what has been an incredible year of innovations, I would like to reflect on a topic which I believe will be one of the defining conversations of 2024: the cost of Generative AI.
Focus On: The Cost of Generative AI and How To Mitigate It
Generative AI is a branch of artificial intelligence that can create new content, such as text, images, music, and code, based on existing data. Generative AI has many potential applications, such as enhancing creativity, improving communication, and generating novel solutions. However, generative AI also comes with a significant cost, both in terms of money and resources.
The financial cost of generative AI
One of the main factors that determines the performance and quality of generative AI models is the amount of data and computation they require. Generative AI models are typically built using foundation models, which are large artificial neural networks that can process massive and diverse sets of unstructured data. For example, ChatGPT, one of the most advanced generative AI models for natural language generation, has 175 billion parameters and was trained on 45 terabytes of text data. According to one estimate, training ChatGPT from scratch would cost about $10 million, and running it for one hour would cost about $25,000. The currently public available version, GPT-4, has 1.6 trillion parameters and could cost more than $100 million to train.
The financial cost of generative AI is not only a barrier for researchers and developers, but also for users and consumers. Generative AI models are often deployed on cloud platforms, which charge fees for accessing and using them. For instance, OpenAI, the company behind ChatGPT, offers a commercial API for accessing its generative AI models, which charges $0.06 per token (a unit of text) for the highest-quality model. This means that generating a paragraph of text with 100 tokens would cost $6, and generating a page of text with 500 tokens would cost $30. While these fees may seem reasonable for some use cases, they can quickly add up for others, as I personally found out when I run Falcon LLM on my private cloud a few months ago.
Hardware is also very, very hard to come by these days. Acquiring the latest, fastest Nvidia GPUs is not just very expensive, but often requires joining a long waiting list.
The energy and environmental cost of generative AI
Another major challenge of generative AI is the energy and environmental impact of its data and computation. Generative AI models consume a lot of electricity, both during training and inference (the process of generating new content from the model). Electricity consumption leads to greenhouse gas emissions, which contribute to climate change and global warming. According to a study by researchers from the University of Massachusetts Amherst, training one large generative AI model for natural language processing can emit as much carbon dioxide as five cars in their lifetimes. Another study by researchers from the University of Pennsylvania and Yale University estimated that generating a single image using a generative AI model can emit as much carbon dioxide as fully charging a smartphone (!!!).
The energy and environmental cost of generative AI is not only a problem for the planet, but also for the people and communities that are affected by it. Generative AI models are often trained and run on data centers. Data centers require a lot of cooling and ventilation, which increase their energy demand and water consumption. Data centers also generate a lot of heat and noise, which can affect the local climate and quality of life. Moreover, data centers are often located in places where electricity is cheap and abundant, but not necessarily clean or renewable. It’s a very similar situation to the one we witnessed during the Bitcoin mining boom, although this time the scale is very different and set to grow.
How to make generative AI more sustainable and responsible
Given the high financial, energy and environmental cost of generative AI, companies need to be careful when shifting from pilots to full scale implementations. Here are some possible steps that can be taken:
Use existing large generative models, don’t generate your own. Unless there is a specific need or benefit, it is better to use existing generative models that have already been trained and optimized, rather than creating new ones from scratch. This can save a lot of time, money and resources, and also avoid unnecessary duplication and waste. For example, instead of training a new generative model for text generation, we can use ChatGPT or other similar models that are available online or through APIs and deploy prompt-engineering and/or custom instructions to direct the output in the right direction.
Fine-tune existing models. If we need to customize or adapt a generative model for a particular domain or task, we can fine-tune it using a smaller and more relevant dataset, rather than retraining it from scratch. Fine-tuning is a process of adjusting the parameters of a pre-trained model to improve its performance on a specific task. Fine-tuning can significantly reduce the data and computation required, and also improve the quality and accuracy of the generated content. For example, instead of training a new generative model for writing blog posts, we can fine-tune ChatGPT using a dataset of blog posts on a specific topic or style. This is by far the best short to mid term solutions for specific use cases, such as content marketing, and can be deployed safely using, for instance, OpenAI models on Azure cloud.
Use energy-conserving computational methods. There are various methods and techniques that can reduce the energy consumption and carbon footprint of generative AI models, such as pruning, quantization, distillation, and sparsification. These methods aim to reduce the size and complexity of the models, while preserving their functionality and performance. For example, pruning is a method of removing unnecessary or redundant parameters from a model, which can reduce its memory and computation requirements. Quantization is a method of reducing the precision or bit-width of the parameters, which can also reduce the memory and computation requirements, as well as the energy consumption. Distillation is a method of transferring the knowledge from a large and complex model to a smaller and simpler one, which can improve the efficiency and speed of the model. Sparsification is a method of making the model more sparse, or less dense, by setting some of the parameters to zero, which can also reduce the memory and computation requirements, as well as the energy consumption. These techniques are essential for companies developing their models in-house.
Use a large model only when it offers significant value. While large generative models can offer impressive and diverse capabilities, they are not always necessary or appropriate for every use case. Sometimes, a smaller or simpler model can achieve the same or even better results, with less cost and impact. For example, if we want to generate a short and simple text, such as a headline or a caption, we may not need to use a large and complex model like ChatGPT, which can generate long and complex texts. We may be able to use a smaller and simpler model, such as BERT or LaMDA, which can generate short and simple texts. Alternatively, we may not need to use a generative model at all, and instead use a rule-based or template-based approach, which can also generate short and simple texts, with minimal data and computation.
Be discerning about when you use generative AI. Generative AI can be a powerful and useful tool, but it is not a magic solution for everything. We should be careful and critical about when and how we use generative AI, and consider the potential benefits and risks of doing so. For example, we should ask ourselves: Do we really need to generate new content, or can we use existing content? What is the purpose and value of generating new content? How will the generated content affect the users? How can we ensure the quality and reliability of the generated content? How can we prevent or mitigate the misuse or abuse of the generated content? These questions are especially true when we think about content marketing and engagement: just because we can generate 10 blog posts a day on a topic does not mean we should. In a world where generative AI will create a lot of vanilla content, ingenuity, quality and human humour will be the winning factors for brands.
Include AI activity in your carbon monitoring. Eventually, we will include our AI activity in our carbon monitoring and reporting, and take actions to reduce our carbon footprint and emissions. For example, we can measure and track the energy consumption and carbon emissions of our generative AI models, and compare them with our goals and targets. We can also implement and report on the steps and strategies that we have taken to make our generative AI more sustainable and responsible, such as the ones mentioned above.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.