How to Reduce Your AI Carbon Footprint
8 Practical Steps for Sustainable Use of Generative AI
The rise of generative AI has sparked meaningful discussions about the environmental impact of ChatGPT, Gemini, Claude and other generative AI technologies. However, these conversations often lack concrete data on actual energy consumption. Google recently disclosed that a median Gemini Apps text prompt uses just 0.24 watt-hours of energy—equivalent to watching TV for less than nine seconds. The company also reports achieving a 33x reduction in energy consumption and a 44x reduction in carbon footprint for these prompts over twelve months, demonstrating that efficiency improvements are accelerating even as capabilities expand. Still, most discussions remain trapped in extremes—either dismissing environmental concerns entirely or suggesting we abandon these transformative tools altogether.
The reality demands a more nuanced approach that acknowledges both documented efficiency gains and the continued need for responsible usage. While these consumption figures are substantially lower than many public estimates, they represent only text generation under optimal conditions. Image generation requires more energy, video generation exponentially more, and training these models in the first place consumes orders of magnitude more resources than any individual query. The generative AI tools transforming how we work and create consume computational resources that vary dramatically based on model size, task complexity, and usage patterns. The good news is that thoughtful users can significantly reduce their environmental impact without sacrificing the benefits these tools provide. Understanding the vast differences between efficient and wasteful AI usage becomes more critical as these tools become ubiquitous in our daily work.
1. Choose the Right-Sized Model for Your Task
The first step toward sustainable AI use involves recognizing that not all models carry the same environmental burden. While headlines focus on massive cloud-based models like GPT-5 or Claude Opus, these computational giants overshadow an ecosystem of smaller, more efficient alternatives. ChatGPT offers "nano" models for simpler tasks, Anthropic provides Claude Haiku, and Gemini offers Flash.
This is how you can access Anthropic’s smallest model, Haiku:
For everyday writing assistance, summarization, basic coding help, and question-answering, small language models like Meta's Llama series, Mistral, or Google's Gemma deliver excellent results while using 10 to 50 times less energy than their heavyweight counterparts. The energy savings ultimately depend on the specific models used.
You can access small language models through cloud services like OpenRouter, which provides dozens of options without technical complexity, or run them locally on recent Apple M-series or higher-end Windows machines using apps like LM Studio or Ollama. Local processing eliminates energy costs from data transmission while keeping your information completely private. Think of model selection like choosing between driving a semi-truck or a compact car for your daily commute. Both vehicles will get you there, but one consumes far less fuel. Reserve the computational heavyweights for truly complex challenges that demand their unique capabilities.
Small models represent a critical intersection of social work values and technological practice, simultaneously reducing environmental impact through lower computational demands and protecting client confidentiality when deployed locally on devices rather than cloud servers. My next article will examine this emerging landscape of compact, efficient models that align with both our professional ethics and sustainability commitments.
2. Establish Mindful Boundaries
Developing a sustainable relationship with AI means recognizing when these tools genuinely add value versus when traditional methods suffice. Spell-checking, basic math, or looking up information rarely requires the computational overhead of generative AI. Before reaching for AI assistance, pause to consider whether the task at hand truly benefits from these capabilities. This boundary-setting isn't about restriction but rather about reserving powerful resources for situations where they deliver meaningful advantages.
3. Understand and Minimize Token Usage
Tokens are the fundamental units that AI models use to process and generate text—think of them as the building blocks of AI language comprehension. Rather than processing complete words, models break text into smaller pieces that can include whole words, parts of words, punctuation, or even individual characters. Common English words, such as "the" or "and," typically form single tokens, while longer words may split into two or three tokens—"understanding" might become "understand" and "ing." Technical terms, proper names, and non-English text often require even more tokens. For instance, "ChatGPT" might be tokenized as "Chat" and "GPT," while a specialized term like "immunohistochemistry" could be broken into five or six separate tokens.
Every token processed or generated consumes computational energy, making the relationship between token count and environmental impact direct and measurable. A response requiring 1,000 tokens uses roughly twice the energy of one requiring 500 tokens. This short article by NVIDIA provides a more thorough description of tokens and their role in AI processing.
You can better understand how ChatGPT tokenizes text by testing its online tokenizer tool. Just enter some text, and a color will represent each token, revealing how the model actually "sees" your input. This visualization often surprises users—seemingly simple sentences can consume dozens of tokens. At the same time, verbose prompts requesting "comprehensive, detailed analysis with multiple examples" can trigger responses consuming thousands of tokens, each one adding to the environmental cost.
Extended reasoning modes, while fascinating for understanding how models process complex problems, exponentially multiply token generation. For routine tasks, these elaborate thinking processes add environmental cost without proportional benefit. Learning to craft concise prompts that yield focused responses, rather than requesting lengthy explorations, can reduce your token footprint significantly. Whenever possible, avoid having generative AI models engage in extended thinking modes when simpler responses will suffice.
4. Improve Your AI Literacy
Developing expertise with generative AI is a strategic approach in its own right. Each regeneration, each reformulated prompt, each abandoned conversation thread consumes additional energy. When you repeatedly ask AI to revise its responses because your initial prompt was unclear, you unnecessarily multiply the computational burden. Prompt engineering—the practice of crafting clear, specific instructions that help AI systems understand precisely what you need—directly reduces your environmental impact. The better you become at prompting, the fewer computational cycles you'll trigger to achieve your desired results.
The most wasteful AI interactions are meandering, stream-of-consciousness conversations that require constant clarification and revision. Instead, adopt a batching approach: collect related tasks and address them in focused sessions with well-prepared prompts. Before initiating any AI interaction, gather all necessary context, define your desired output clearly, and craft comprehensive instructions offline. This preparation dramatically reduces the back-and-forth exchanges that multiply energy consumption. A single well-crafted prompt that yields the right result immediately uses far less energy than a dozen attempts at refinement.
Here is a recent guide on prompt engineering skills offered by DataCamp that can help you develop this expertise.
5. Support Sustainable Providers
The choice of AI provider extends beyond features and pricing to encompass environmental responsibility. Some companies power their data centers entirely with renewable energy, while others invest heavily in carbon offset programs or energy-efficient hardware development. Microsoft has committed to being carbon negative by 2030, while Google aims to run on carbon-free energy 24/7 by the same year. Smaller providers like Mistral actively contribute to global environmental standards for AI development.
These differences matter. A provider running on coal-generated electricity produces significantly more emissions per query than one powered by solar or wind. By researching and supporting providers who prioritize sustainability, you send clear market signals about consumer values, potentially influencing industry-wide practices. Consider checking providers' sustainability reports before committing to subscriptions, asking about their renewable energy usage when evaluating enterprise solutions, and prioritizing companies that publish transparent environmental impact data. Your subscription choices become votes for the kind of technological future you want to see.
Here are two important reports from Google and Mistral.
6. Contextualize Your AI Usage Against Other Digital Activities
Understanding AI's environmental impact requires context: How does it compare to other digital activities we rarely question? Google's recent disclosure that a median Gemini AI text query consumes just 0.24 watt-hours provides a valuable benchmark; however, direct comparisons to other activities should be approached cautiously, as energy consumption varies widely based on specific devices and usage patterns.
The contrasts become more striking with entertainment consumption. Streaming video typically consumes 70-80 watts of power across all components (TV/device, router, and data centers). Watching 90 minutes of Netflix at 77 watts would use approximately 116 watt-hours of energy, which is equivalent to about 480 Gemini text queries. For illustration, if a small local model (e.g., Llama 3.1 8B) uses roughly one-twentieth the energy of Gemini (approximately 0.012 watt-hours per query), those same 116 watt-hours could theoretically power nearly 10,000 queries. However, actual consumption varies significantly based on hardware and implementation.
To put this in a broader perspective, binge-watching the complete Game of Thrones series (approximately 70 hours across 8 seasons) at 77 watts would consume about 5,390 watt-hours. This amount of energy could power roughly 22,500 Gemini queries or 450,000 queries on a small local model, assuming these efficiency estimates hold. This perspective shift doesn't minimize AI's environmental impact; instead, it helps calibrate our concerns appropriately.
These comparisons should be understood with important caveats: video streaming energy use is well-established and relatively consistent, while AI energy consumption varies dramatically based on task complexity, model size, and whether you're generating text, images, or video. Additionally, Google's 0.24 watt-hour figure represents optimal cloud infrastructure performance, and actual consumption may vary based on provider efficiency and specific use cases. The key insight is that text-based AI, when used purposefully for work and creativity rather than endless experimentation, can be remarkably efficient compared to many everyday digital activities. However, image and video generation require orders of magnitude more energy, making task selection a critical factor in sustainable AI use.
7. Choose Generation Types Wisely
Text generation requires orders of magnitude less computational power than image creation, which itself is significantly less demanding than video generation's resource requirements. While Google has documented that a median Gemini text query uses 0.24 watt-hours, comparable data for image and video generation from major providers remains limited.
Research estimates suggest the energy hierarchy follows this general pattern:
Text generation: Minimal energy (Google's Gemini: 0.24 watt-hours)
Image generation: Substantially higher (estimates range from 2-10x more than text, depending on resolution and model)
Video generation: Exponentially more intensive (potentially hundreds to thousands of times more than text, scaling with length and quality)
The stark differences stem from computational complexity. Text generation processes sequential tokens, while image generation must compute millions of pixel relationships, and video generation multiplies that complexity across numerous frames. Though specific consumption figures for tools like DALL-E, Midjourney, or Sora aren't publicly available, the underlying computational requirements make their relative energy intensity clear.
When visual content isn't essential, opting for text-based outputs dramatically reduces energy consumption. If images are necessary, specificity in your prompts is critical because each iteration seeking that "perfect" visual multiplies environmental impact. Consider whether a detailed text description might serve your purpose as effectively as a generated image, especially for internal documentation or communication where visual polish may be less critical than clarity.
8. Create Reusable Resources
When AI generates beneficial content, consider how to preserve and share these outputs to avoid the environmental cost of regenerating similar content repeatedly. Each time you or a colleague re-prompts AI for comparable results, you're triggering unnecessary computational cycles that consume energy and contribute to carbon emissions.
Building personal or team libraries of AI-generated resources, documenting successful prompts for routine tasks, and sharing effective outputs with colleagues transforms individual AI usage into collective environmental responsibility. Think of it as recycling at the computational level. A single well-crafted prompt that gets shared across a team of ten people eliminates nine redundant processing cycles, multiplying your environmental savings without sacrificing productivity.
Moving Forward Consciously
The path toward sustainable AI usage doesn't require abandoning these powerful tools but rather approaching them with the same environmental consciousness we bring to other aspects of modern life. While this guide focuses on the choices within our direct control (e.g., model selection, prompt efficiency, and usage patterns), the most significant environmental costs of AI lie beyond the influence of individual users.
Training large language models can consume millions of times more energy than individual queries, requiring massive computational resources over weeks or months. Additionally, the environmental footprint includes water usage for cooling data centers, rare earth mining for hardware components, and eventual equipment disposal. However, since technology companies bear these infrastructure and training costs that remain outside individual users' control, our focus remains on optimizing what we can influence: how we utilize these already-trained models in our daily work.
Through informed model selection, mindful usage patterns, skill development, and strategic choices about when and how we engage with AI, we can harness the benefits of these technologies while minimizing their environmental impact. Every conscious decision—choosing a smaller model for simple tasks, crafting precise prompts, or sharing successful outputs with colleagues—represents meaningful action within our sphere of influence. The goal isn't perfection but progress, making each deliberate choice a contribution to a more sustainable relationship between human creativity and artificial intelligence.
This framework for sustainable AI usage will evolve as technology advances and our understanding deepens. What remains constant is the principle that environmental responsibility and technological innovation need not be mutually exclusive. By adopting these strategies, you join a growing community of users who demonstrate that we can harness AI's transformative potential while minimizing our environmental footprint, focusing our efforts where they can make the most difference.



