Published October 1, 2026 | Version 1.0.1

"So, What's the Carbon Cost of Using AI?" A Bounded Estimation Framework for Calculating the Per-Token Carbon Intensity of Large Language Model Inference Under Limited Disclosure, Using the Worked Example of OpenAI's GPT-5 Model Family

Authors/Creators

  • 1. InferenceCarbon Ltd

Description

So, what is the carbon cost of using AI? No leading provider tells us: OpenAI, Anthropic and others publish no per-model, per-query energy or carbon figures. This means that nothing is disclosed about the operational carbon footprint of inference at the one unit where a user decides – a particular prompt, model and effort setting. We present a bounded estimation framework that answers the question from public and collected data alone. It triangulates an anchor (a published energy figure for one model) across three classes of public source, scales between models by observed throughput, applies query-length-conditional reasoning multipliers measured by our token-level campaign, corrects for the rise in per-GPU power between hardware generations, reports location-based carbon (what the local grid emitted) and market-based carbon (net of the provider's clean-energy contracts) in parallel under the GHG Protocol Scope 2 Guidance, and compounds residual uncertainty into a single multiplicative envelope (a range around each figure).

Applied to OpenAI's GPT-5.x family, the framework estimates a fleet intensity of 353 gCO₂e/kWh location-based (319 – 500 across routing scenarios) and 52 gCO₂e/kWh market-based (47 – 65, rising to 159 in a particular scenario.) Per-model central estimates – per 1,000 tokens, about 750 words – span roughly 14×, from 1.07 gCO₂e/1,000 tokens location-based (0.16 market-based) for GPT-5.4-nano with reasoning off to 14.80 (2.16) for GPT-5 at high effort – within envelopes of roughly 10 – 19× top-to-bottom, median ~11×. All headline figures carry Low confidence (IPCC AR5 – the evidence is limited): a single line of audited per-model disclosure would collapse most of the envelope, and we invite it.

What should a user do? A general user need not worry – twenty short prompts a day for a year is roughly 4 kg CO₂e, about 0.04% of the UK's average per-person footprint – however framing the task clearly and asking for shorter answers costs little. A power user, at about 122 kg a year (a short-haul plane flight), should choose the smallest capable model and the lowest effort setting the task allows – on GPT-5 to 5.4 the effort setting is the larger lever, from GPT-5.5 onward model choice is. An application designer holds the largest lever: at a million queries a month, routing from a flagship model at high effort to a mini model at medium could save on the order of 38 tCO₂e a year, the equivalent of just under four business-class return flights from Dallas to Delhi.

Notes

Version 1.0.1 of the paper text (1 October 2026). It revises the text of v1.0.0 (25 September 2026) for clarity; every table is unchanged. The tables are computed from the v1.0.0 data and code release, archived separately at https://doi.org/10.5281/zenodo.22962476. Not yet peer-reviewed.

Files

InferenceCarbon_v1_0_1.pdf

Files (1.7 MB)

Name Size Download all
md5:d80d28f7bb8c0a0944f6c7e21e113b35
1.7 MB Preview Download

Additional details

Related works