Mastodon Feed: Post

Mastodon Feed

Boosted by pixxl:
jonny@neuromatch.social ("jonny (nonvenomous)") wrote:

@kepeken the numbers here are astronomical - 69 billion tokens in a month, or, for a rough average, 50-60 billion words worth of text on the fable tokenizer. there are 5.2 billion words on english wikipedia, or, it's about 92,000 copies of War and Peace back to back.

Nobody publishes numbers on this, but like, back of the envelope, think about the energy cost alone:
Fable is something like a 2-5 trillion parameter model, google estimated their ~200B model at ~0.24Wh/prompt last year. that's probably low, since it was a google-authored study, but we'll go with it. So assume linear energy scaling per parameter (i would bet it's actually supralinear), that's ~15x the size, or 3.6Wh/prompt.

google didn't say what the median token count in a "prompt" was, but wild guess, say the median prompt + response that google would be measuring is something like 3 paragraphs, or 300 tokens. Then you have 3.6/300=0.012Wh/token.

So then the astronomical numbers part: 0.012Wh/token * 69 billion tokens = 828 million Wh, or 828 thousand kWh. The average electricity price in the US in 2025 was $0.136 per kWh.

So, that means, back of the envelope, estimating conservatively, the energy should be in the ballpark of 828k * $0.136 or $112,000, or like 1.3x the API cost. We would have had to be off by a factor of 1.3 to break even on the energy costs alone - note that the google estimates were just of inference not including training. Models could have gotten way more efficient, but a number of the guessed parameters could have gone the other way too: energy more expensive, fable is larger than 2T, energy scales supralinearly, eg. And the conventional wisdom is that inference is orders of magnitude cheaper than training, so we need to be way better than breaking even on inference to make up for all the capex spend.

So it might seem like $80k for a months worth of tokens suggests that the API prices are inflated, but a) this person is generating an absolutely preposterous quantity of text, and b) the thing that is being done is so much more inefficient than you can possibly imagine, it's like trying to think about how much bigger the sun is than your car.

editing: swapped 69->96 billion tokens initially, it's 69 billion.