ZUMENTIQ

AI & digital transformation

What does an AI query really cost? Google, ChatGPT and the data centres behind them

Electricity is only part of the AI bill. Understand search, ChatGPT and data centre costs with transparent examples for Swiss SMEs.

The wrong question: “How much more expensive is AI than Google?”

A Google search costs nothing. Neither does a question to ChatGPT, at least when no individual charge appears. Meanwhile, billions are being invested in data centres. The explanation is straightforward: the price visible to the user tells us little about the infrastructure behind it. Turning that into a fixed claim such as “AI uses ten times more electricity” skips the questions that matter.

Are we finding a phone number, drafting an email or asking an agent to compare twenty documents? These are different services. For a Swiss SME, the goal is not the cheapest query but an acceptable, verified result at a sensible total cost. Looking under the bonnet still helps: it reveals where costs arise and which decisions actually make a difference.

Sources checked on 18 September 2026. The following sections distinguish provider disclosures, research estimates and our own illustrative calculations. Exact, current fully allocated costs per Google or ChatGPT query are not publicly available.

Three different bills

Customer prices are subscriptions, licences or usage-based API charges. A flat subscription is spread across your actual usage; it does not establish a technical cost per request. Ad-funded search is not free to operate either. The bill is simply paid differently.

Operator costs include hardware, depreciation, financing, staff, networking, buildings, maintenance and energy. Depending on the scope, they also include research, training and product development. The incremental cost of another request is not the same as its share of total costs. A low selling price therefore proves neither low costs nor profitability.

Electricity costs are energy consumed multiplied by a tariff. Consider a deliberately simple example: 0.3 Wh equals 0.0003 kWh. At an assumed CHF 0.15–0.30/kWh, that costs CHF 0.000045–0.000090, or 0.0045–0.009 Swiss centimes. This is neither a ChatGPT price nor its verified cost structure. It simply illustrates that a tiny electricity bill can coexist with substantial hardware and development expenses.

What the comparison with search actually tells us

Google reported approximately 0.3 Wh per search in 2009, including allocated work before the request, such as building the index. That is a historical reference, not a measurement of Google Search in 2026. Search with an AI summary also combines retrieval and generation. “Google” and “AI” are no longer mutually exclusive categories.

Request typePublished figureWhat it does not establish
Traditional Google search0.3 Wh; Google, 2009A current comparison baseline
Gemini text responseMedian 0.24 Wh; May 2025 production dataA figure for every AI task
ChatGPT requestAverage 0.34 Wh; Sam Altman, June 2025A fully documented measurement boundary
Search with an AI answerNo robust universal figure in the sources reviewedEquivalence with Gemini Apps
Long reasoning taskHighly dependent on tokens and executionA fixed multiplier over traditional search

The Gemini study includes active accelerators, CPU and memory, reserved idle capacity and data centre overhead. Its energy figure concerns inference, not training or a complete lifecycle including hardware production and the user's device. A median is not a mean. Altman's disclosure lacks sufficient methodology for a like-for-like comparison. These numbers do not support a league table.

Why AI requires so much computation

During training, a model learns from large datasets: it makes predictions, evaluates errors and updates a very large number of parameters over many iterations. This development phase is not repeated for every user question. Economically, however, its costs must be recovered through subsequent usage or other revenue.

During inference, the trained model uses its parameters to generate a new response. Language is divided into tokens: small text units, not necessarily whole words. The system first processes the input context, then generates its answer step by step. A long conversation history and a lengthy output require more work than a short, narrowly defined task.

GPUs and other AI accelerators are designed to perform many operations in parallel. They process the large numerical arrays used by models more efficiently than an ordinary office processor. They still need fast memory, data connections and electricity. Large models can span multiple chips. Moving data and keeping capacity available are part of the workload, not just the calculation itself.

Reasoning, response length and utilisation change the economics

Reasoning systems may generate additional tokens and intermediate steps before producing the visible answer. Agents call tools, check results and initiate further model requests. One user question can therefore trigger an entire processing chain. The length of the final response does not reveal all the computation involved.

A study by Oviedo and colleagues, revised on 9 June 2026, models large systems on H100 nodes. Typical queries have an estimated median of 0.31 Wh, with the middle 50% spanning 0.16–0.60 Wh. A scenario with 15 times as many tokens reaches 3.91 Wh, with an interquartile range of 2.15–7.05 Wh. These are estimates under defined assumptions, not measurements of all ChatGPT users; the ranges are not universal minimum and maximum values.

Utilisation matters too. Processing multiple requests together spreads infrastructure overhead across more results. Capacity reserved for traffic peaks and fast responses costs money even when it is doing little work. That is why a GPU's maximum wattage alone cannot establish credible energy consumption per query.

What an API request can cost

Anthropic's price list provides a verifiable example: Claude Sonnet 4.6 has standard rates of USD 3 per million input tokens and USD 15 per million output tokens, checked on 18 September 2026. This is a specific example, not a recommendation of the newest or cheapest model. Other plans, regions and additional services may differ.

With 2,000 input and 500 output tokens, our calculation is 0.006 + 0.0075 = USD 0.0135, or 1.35 US cents. With 20,000 input and 5,000 billable output tokens, the result is USD 0.135. These examples exclude caching discounts, batch processing, tools and taxes. Both are selling prices, not electricity measurements.

When budgeting, calculate the entire workflow: how many model calls, retries and reviews are needed per completed task? A small token price can become material with long documents and unrestricted agent loops. Conversely, a more expensive model may be more economical if it needs significantly less correction.

What a data centre costs: the building is not the computer

The Turner & Townsend 2025–2026 construction index gives Zurich a benchmark of approximately USD 14.2 per watt of IT capacity. Multiplying by 30 MW gives roughly USD 426 million. The benchmark concerns air-cooled hyperscale facilities, typically 30–50 MW. It is neither a Swiss quotation nor the cost of a fully equipped AI data centre.

The methodology includes construction and mechanical and electrical equipment. It excludes items such as active IT, land, external utility works and professional fees. Transformers, backup power and cooling belong to the infrastructure. Servers, GPUs, storage and data networks are another bill.

IREN's announcement of 3 November 2025 illustrates its scale: approximately USD 5.8 billion for GPUs and ancillary equipment associated with a planned 200 MW IT deployment. This includes servers, networking, software and deployment services. The separate USD 9.7 billion figure is a five-year customer contract, not construction expenditure. The US project cannot simply be extrapolated to Zurich.

A transparent investment and operating example

The following hypothetical 10 MW scenario illustrates the scale. It is our own calculation, not a market quotation. The ranges are chosen sensitivity assumptions, not statistical confidence intervals. Their purpose is to expose the inputs a serious budget would require.

ItemExplicit assumptionResult
Building and technical infrastructure10 MW × CHF 10–15 million/MWCHF 100–150 million
IT including GPU systemsSeparate assumed budgetCHF 150–300 million
Investment subtotalExcludes land, external site connections, financing, taxes and other project costsCHF 250–450 million
Annual electricity70% average electrical IT load, PUE 1.20, 8,760 hours73.584 GWh
Annual electricity billCHF 0.10–0.20/kWh assumed effective tariffApproximately CHF 7.4–14.7 million

Adding simplified straight-line depreciation over 15 years for all infrastructure and four years for IT, with no residual value, produces approximately CHF 44–85 million annually. Actual components have different useful lives. Combined with electricity, that is roughly CHF 52–100 million a year, before staff, maintenance, spare parts, licences, insurance and capital costs.

This is not an estimate of OpenAI's or Google's operating costs. Without productive throughput, workload mix and lifetime assumptions, it cannot establish a cost per query either. Those missing inputs explain why apparently precise calculations online are often misleading.

MW, kWh and PUE: getting the units right

MW measures power: the rate at which electricity can be or is being consumed. kWh measures energy: power multiplied by time. One megawatt sustained for one hour equals 1,000 kWh; one GWh equals one million kWh. A campus capacity figure is therefore not an annual consumption figure.

In our example, 10 MW × 70% gives 7 MW average electrical IT load. This is not 70% GPU compute utilisation; they are different measures. PUE 1.20 means the data centre consumes 1.20 kWh in total for each 1 kWh used by IT. Therefore: 7 × 1.20 × 8,760 = 73,584 MWh. The extra 20% is relative to IT energy; it represents approximately 16.7% of total consumption.

Cooling, power conversion and other facility systems account for this overhead. PUE alone says nothing about the usefulness of a computation or its carbon footprint. Water consumption and manufacturing impacts require separate boundaries. For a tangible comparison, 1 Wh powers a 10-watt LED for six minutes. That is energy equivalence, not a complete environmental assessment.

Small individual requests, large collective impact

In its 2026 update, the IEA puts global data centre electricity consumption in 2025 at approximately 485 TWh and projects around 950 TWh for 2030. The second figure is a scenario, not an observed future. Both cover data centres overall, not just generative AI.

The apparent contradiction disappears when volume enters the picture: billions of small operations, more video, longer contexts and continuous automated loops add up. More efficient chips reduce energy per task; additional demand can outweigh those gains. For an SME, neither blanket avoidance nor indiscriminate automation is the right response. Remove unnecessary processing and deliberately budget for processing that creates value.

Five practical decisions for Swiss SMEs

  • Separate tasks: Use search for a known address. A smaller model may suffice for summarising or drafting. Reserve reasoning for work where extra checking demonstrably helps.
  • Use a real test set: Compare twenty representative tasks. Record model charges, turnaround time, errors and rework. Do not upload confidential material to unapproved services.
  • Limit context: Send relevant passages instead of entire archives. Specify the desired response length and cache reusable information where appropriate.
  • Put boundaries on automation: Set budgets, iteration limits, timeouts and human approval points. A retry loop must not continue indefinitely.
  • Compare complete operating costs: For self-hosting, include low utilisation, administration and hardware replacement. For cloud services, include integration, data handling and oversight.

My recommendation is to start with one frequent, low-risk task and a clear acceptance criterion. Scale only when quality and total costs make sense. Saving a fraction of a cent on the model achieves little if somebody then spends another ten minutes correcting the result.

Conclusion: not every prompt needs a power station

A short text request can consume very little electricity. That does not make AI infrastructure inexpensive. The billions fund a combination of capacity, hardware and reliable operations. Conversely, a large data centre does not make every individual use wasteful. Distinguishing cost categories, measurement boundaries and business value leads to better decisions than any blanket “AI versus Google” multiplier.

Which AI use case makes financial sense for your SME?

Let us assess the task, quality and total costs together, before an experiment becomes a recurring budget.

Book a free initial consultation