Hi, Shrewd!        Login  
Shrewd'm.com 
A merry & shrewd investing community
Best Of GOOGBest OfAll BoardsThe Shrewd’m WeeklyLearn to InvestHow to Become Shrewd
Search
Shrewd'm.com Merry shrewd investors
Search
Best Of GOOGBest OfAll BoardsThe Shrewd’m WeeklyLearn to InvestHow to Become Shrewd


The week's question
In December 2024, in the thread "Re: BRK: Why Not XOM?", BreckHutHigh asked the members: "What about the long road trips with kids?" This week it is put to everyone again. The button below opens the small thread re-asking it - read what others have said so far, then give your own answer as an ordinary reply.
Answer this questionContinue to Shrewd'mThis note won't appear again
Stocks A to Z / Stocks G / Alphabet (GOOG)
Unthreaded | Threaded | Whole Thread (2) |
Author: mungofitch 🐝🐝🐝 GOLD
SHREWD
  😊 😞

Number: of 442 
Subject: Cost of doing business
Date: 08/26/26 7:40 AM
Post New | Post Reply | Report Post | Recommend It!
No. of Recommendations: 13
Across the corporation, Alphabet is doing a whole lot of things on the LLM front. Products, chips, lending, borrowing, leasing, building, lobbying, you name it.

But I wanted to focus on one very specific thing: old fashioned Google search, and changes to the cost of providing it.

The AI overviews are an interesting phenomenon. They're quite good, especially the ability to ask follow on questions, but I was worried it might be a "red queen" phenomenon: they cost a lot of money to provide, but don't really increase search revenue directly (hard when you're close to a monopoly already), so it's now just a higher cost of business--neither Google nor any putative competitor is any better off, but users now expect it. Clearly there is a lot more CPU power required to provide the AI summaries than the old rank algorithm, and that costs money. We keep hearing about exploding costs of tokens (demand rising faster than prices are falling), so is that an internal problem for Alphabet?

So I thought I'd ask an LLM what it thought about this issue. Not Gemini, as it might be in a conflict of interest situation : ) This is from Claude Opus 5.

What follows is synthetically generated, except my prompts in italics.

=========================

Google has recently added AI summaries to the default search results. A fraction of these lead to follow on questions to the LLM. Plus, of course, the original search results list is provided. Estimating a weighted average among these tasks, what is the factor of increased computing demand per query session, relative to the baseline before the introduction of AI summaries?

Short answer: roughly 2–3× per query session, with a defensible range of about 1.3× to 8×. Here's the arithmetic behind that.

The model. Take one pre-AI-Overview session as 1 unit. Three things now happen:
Component                                           Probability              Cost (units)      Contribution
Classic retrieval + ranking (still runs in full) 1.0 1 1.00
AI Overview generation ~0.3 ~3 0.90
Follow-on LLM turns (AI Mode, "dive deeper") 0.3 × 0.10 × 1.5 turns ~8 0.36
Total ~2.3×

Where each number comes from.
Trigger rate. This is the most contested input. Trigger rates vary between studies from roughly 13% to 60% depending on keyword sample and method. Semrush tracked AI Overview prevalence climbing from about 6.5% in January 2025 to a peak near 24.6% in July before settling around 15.7% by November, while BrightEdge-derived trackers put the rate at roughly 48% of queries as of early 2026. Weighted for the fact that high-trigger categories (comparison, question-form, health) are also the categories where users linger, ~0.3 is a reasonable central value.
Cost ratio per overview. Google's own production telemetry puts the median Gemini Apps text prompt at 0.24 Wh, against a search baseline that is genuinely unknown — the widely-quoted 0.3 Wh comes from a 2009 Google estimate with no update since, and independent 2024 estimates put modern search nearer 0.04 Wh. So the ratio of one LLM generation to one search is somewhere between ~1× and ~6×. An AI Overview should sit below a median Gemini prompt (smaller distilled model, short output, and overviews for popular queries are cached and amortised across many sessions) but above it on input tokens, since it ingests retrieved passages. Call it 3×.
Follow-ons. These are rarer than intuition suggests but individually expensive, because AI Mode does query fan-out — issuing multiple background searches per turn — so it multiplies both the retrieval and the generation cost. Sessions are also longer: AI Overview sessions average about 21 seconds of user time per query versus 49 seconds for AI Mode.

Three caveats that matter more than the point estimate:
1. The denominator is the weak link. Nobody outside Google knows what a 2026 keyword search costs, and modern ranking already runs transformer models. If search is 0.04 Wh, the multiplier goes toward the top of my range; if it's 0.3 Wh, toward the bottom.
2. Per-session multipliers are falling fast. Google reports the median Gemini prompt used 33 times more energy in May 2024 than in May 2025 — so the ratio is a moving target, dropping per unit even as coverage grows.
3. Total demand rises faster than per-session demand, because of induced usage. Google has said AI Overviews drive about 10% more search usage for the query types where they appear. Multiply your ~2.3× by that and fleet-wide demand is up ~2.5×.
One thing to discount: the "AI search uses 10× a standard search" headline. That traces back to a 2023 remark by Alphabet chairman John Hennessy, converted into a ~3 Wh per-interaction estimate and divided by the 2009 search figure — an outdated numerator over an outdated denominator.

One possible sign that the computing cost is onerously costly and/or so demanding as to be capacity constrained is that the rollout of AI overviews has been very slow...both by geography with gradually staged rollouts by country,and by frequency, with only a subset of queries likely to benefit from an AI summary likely to get one. In short, there are reasons to believe that the incremental burden is at the high end, and or being minimized in publicly disclosed figures.

It's a reasonable hypothesis, and it's half right — but the specific evidence you cite points the other way, and it's worth separating the two claims.

Where you're right: cost was clearly a first-order constraint. You don't repeatedly brag about a cost reduction unless it was a problem. Pichai told the Q3 2024 call that in eighteen months Google had reduced costs by more than 90% for AI Overview queries through hardware and engineering breakthroughs, while doubling the size of the custom Gemini model, and the campaign hasn't stopped: after moving AI Overviews and AI Mode to Gemini 3, they reported cutting the cost of core AI responses by a further 30%. Backing that: Alphabet's capex ran $136 billion over the trailing twelve months with another notable increase expected in 2027. So yes — the launch-era multiplier was almost certainly enormous. A 90% reduction implies the mid-2023 AI Overview cost something like 10× what it cost by late 2024, which would put the original figure in the 20–50× range against a plain search. That is exactly the regime in which you stage a rollout.

Where the inference breaks: the selectivity pattern is wrong for capacity rationing. If scarce inference capacity were binding, you would spend it on your most valuable queries. Google does the opposite. Healthcare triggers at 88%, education 83%, B2B tech 82% — while e-commerce fell to roughly 4%. Informational queries trigger at 36%, commercial at 8%, transactional at 5%. Transactional queries are where the ad revenue is. A capacity-constrained operator rations the cheap, low-value queries first; Google is suppressing overviews precisely where revenue per query is highest. That's ad-cannibalisation management and quality gating, not compute rationing.

The geography does the same thing. Trigger rates are highest in Indonesia at 37.2%, then the Philippines and Mexico at 29.1%, India 26.8%, Nigeria 26.4% — with the US 13th at 20.5%. Those are the lowest ARPU markets on the list. No one short of serving capacity allocates it there first. And the geographic staging itself was not slow: AI Overviews went to more than a hundred new countries and territories in a single week in October 2024, reaching over a billion monthly users. The pacing correlates far better with language quality, the mid-2024 output fiasco, and EU regulatory and publisher exposure than with FLOPs.
There's also a mechanism that cuts your estimate rather than raising it: trigger rates are measured by rank trackers observing SERPs, but overviews for repeated queries are cached and served without regeneration. Observed presence overstates actual generations, possibly substantially, since query distributions are heavily head-weighted.

On disclosure minimisation — you have a real point, but a narrow one. The 0.24 Wh figure is a median on a distribution that is severely right-skewed by long outputs and reasoning; the mean is certainly higher and Google chose not to publish it, nor the query volume needed to reconstruct totals. It also excludes training amortisation. And the comparison is flattered by a denominator Google has declined to update since 2009. Those are fair criticisms of the framing. But the 33× year-over-year improvement is mechanistically credible — distillation to Flash-Lite-class routing models, speculative decoding, quantisation, batching, and TPU generational gains each contribute several-fold, and they compound.

The synthesis, which I [sic] think is the actually interesting part: a binding constraint and a low per-query multiplier are not in tension — they're causally linked. A constrained operator optimises until marginal cost fits marginal revenue, then expands coverage. The optimisation campaign is the evidence of how painful it was, and also the reason the ratio settled where it did. Notably, Pichai has said latency is more of a constraint than cost per query — which suggests the binding limit is time-to-first-token and serving capacity, not energy.
So I'd revise the composition rather than the level: hold the central estimate near 2.5×, but widen the upper tail, because AI Mode's query fan-out — not the overview itself — is now the real cost centre, and it's the piece with the least public disclosure.

===========================

Possible bottom line (from the human here): given the stonking margins of the pre-LLM search business, an increase of 2-3x in infrastructure costs to the COGS in that specific division isn't a threat to the company's business and probably won't hurt profitability in any meaningful way, given the number of more important moving parts to the business.

Jim

Post New | Post Reply | Report Post | Recommend It!
Print the post
Members reply directly to mungofitch here — and replies get answered. Reading is free; so is joining the conversation. Join Shrewd'm »
This community has written 442 posts about Alphabet. The article-length ones it recommended most:
Judge's Ruling · 24 recs · 2025
Alphabet's Century Long Bet · 23 recs · 2026
Will the Sexy Six Strut or Stumble? · 19 recs · 2025
2Q Summary 2026 · 18 recs · 2026
Up about 50% this year · 17 recs · 2023
Unthreaded | Threaded | Whole Thread (2) |


Announcements
Alphabet FAQ
Contact Shrewd'm
Contact the developer of these message boards.

Best Of GOOG | Best Of | Favourites & Replies | All Boards | Followed Shrewds | Open Questions | Moving a community