A startup could reach a two-trillion-dollar valuation in under six years, and half the internet is arguing about whether that's a bubble. Both camps are answering the wrong question. AI isn't weightless software. It's a physical production system, and the number that outlives the hype cycle is cognitive work per kilowatt-hour.
tsukumo
Short version: the argument about whether AI is a bubble is loud, and it is aimed at the wrong target. Both sides treat AI as software economics, which quietly assumes the thing is nearly free to run. It isn't. AI is a physical production system: capital buys electricity, electricity buys compute, compute produces tokens, tokens do useful work. A token costs kilowatt-hours. Once you see it that way, the durable question stops being "bubble or not" and becomes: how efficiently does capital and electricity convert into useful cognitive output? That is the number I'd watch, and it is the one the shouting drowns out.
By one investor scenario, an AI company could carry a valuation near two trillion dollars about 5.8 years after it was founded. For scale: Alphabet took roughly 23 years to cross that mark, Amazon about 30, NVIDIA about 31, Apple 44.4, Microsoft 46.2. The bar that took the old giants a working career took the new one less than a college-to-mid-career gap.
Years from founding to a first $2T valuation. The 5.8-year bar is a projected investor scenario for a potential October 2026 IPO, not an achieved public-market valuation; the others are realized (Alphabet Nov 2021, Amazon Jun 2024, Apple Aug 2020, Microsoft Jun 2021). Sources, FT and Reuters; founding-to-$2T years derived.
I need to be exact about that first bar, because the whole essay is built on being honest with numbers. The 5.8-year figure is a projected investor scenario tied to a potential October 2026 IPO. It is not a valuation the public markets have set and cleared. Treat it as a bet being priced, not a fact on a ticker. Even discounted for that, the velocity is the story: the curve from founding to a two-trillion valuation has compressed by roughly an order of magnitude in a single generation of companies.
Look at what the market is actually reacting to, because it isn't only a revenue line. A two-trillion price on a company a few years old is a bet on a rate of change, not on this year's cash flow. The thing that re-rated is the belief about how much economic work one unit of this technology can do, and how fast that number is still climbing. When the perceived slope of "useful output per dollar of compute" jumps, the terminal value jumps with it, and it jumps faster than any single quarter's numbers can explain. That's what a compressed founding-to-$2T curve is: a market repricing a slope.
So the internet did the internet thing. One camp says the numbers are detached from any plausible cash flow and the correction is coming. The other says this is the biggest platform shift since the smartphone and the skeptics will look like the people who called the web a fad. Both camps are sure. Both are arguing about a multiple.
Here is my claim, stated up front so you can decide whether to keep reading: they are answering the wrong question, and they share the reason why.
Notice what the bubble debate takes for granted. It argues about the price of AI while assuming it knows the nature of AI, and the assumed nature is: software.
Software economics has a signature shape. You pay a large fixed cost to build the thing, then the marginal cost of one more user rounds to zero. Distribution is nearly free. The dot-com comparison rides entirely on that shape, and so does the "it's just hype" dismissal, and so does the "marginal cost of intelligence goes to zero" techno-optimism. All three are the same mental model wearing different moods.
That model is what's wrong. Not the optimism, not the pessimism. The model.
Because a token is not a free copy of a file. Producing it consumes electricity, occupies a GPU for a slice of time, and throws off heat that something has to move. Serve a million users and you are not duplicating a static asset a million times at zero cost. You are running a million loads through a physical plant that draws power the entire time. The marginal cost of the next unit of AI output is not approximately zero. It is a metered quantity of energy.
When your cost model is wrong by that much, every valuation argument built on top of it is arguing about the wrong axis. The dot-com boom mispriced companies that were, underneath, genuinely software: the correction came, and then the survivors compounded on near-zero marginal cost for two decades. AI does not get that second act for free, because its marginal cost never collapses to zero. It bottoms out at the price of a kilowatt-hour.
So the useful move is not to pick a side in the bubble fight. It is to change what we are measuring.
Here is the chain I actually keep in my head when I reason about this, walked once, end to end:
capital → electricity → compute → tokens → useful work → human effort displaced → economic output.
Capital buys the two things that matter: hardware and the power to run it. Electricity is converted, through compute, into tokens. Tokens, when the system is any good, get turned into useful work: a merged change, a drafted contract, a resolved ticket, a structured decision. That useful work displaces some quantity of human effort, and that displacement is what shows up, in aggregate, as economic output.
This is a factory. Not metaphorically. Structurally. It has a raw input with a spot price (electricity), a conversion step with a throughput limit (compute), and an output that has to be worth more than the input for the thing to make sense. The reason it doesn't feel like a factory is that the product is abstract and the plant is a warehouse of silicon in a county you have never visited. But the accounting is the accounting of industry, not of software.
Make it concrete with one unit of work. Say an agent resolves a bug: reads the failing test, finds the cause, writes the fix, checks it. To do that it runs some number of forward passes through a model, and each pass draws power off the GPUs for the milliseconds it takes. That run consumed a measurable quantity of energy and a slice of a machine's time, and it produced a result that was worth some fraction of an engineer's hour. Every link in the chain is present in that one task: the capital that bought the GPU, the electricity it drew, the compute it occupied, the tokens it emitted, the useful work those tokens amounted to. Run that a billion times a day and you have not copied a file a billion times at zero cost. You have run a billion loads through a plant that was drawing power the whole time. The unit economics of that are industrial, not editorial. There is a real input cost on every single unit, and it never goes to zero no matter how many units you make.
That is the part the software framing cannot see, because in software the second unit really is almost free. Here it isn't. Here the second unit costs about what the first one did, minus whatever efficiency you've won at the hardware layer. Which is exactly why the efficiency layer turns out to be the whole game, but I'm getting ahead of the argument.
Now, one link in that chain needs a hard fence around it, and I want to build it before anyone misreads the sentence.
I hold that distinction because it is easy to sound either naive or cynical here, and both are wrong. Naive: pretend nothing is displaced, that it's all "productivity." Cynical: pitch the displacement as the feature. The accurate version is drier than either. Displacement is the economic variable. Augmentation is the engineering stance. Keep them in separate columns.
With that fence up, the interesting question becomes obvious. If AI is a factory whose input is electricity, then the constraint that governs it is not clever code. It's power.
Every factory has a binding constraint, the input that runs out first. For AI, at the scale capital is now betting on, that input is electricity.
Global data-center electricity demand, about 415 TWh in 2024 to roughly 945 TWh by 2030, close to a doubling in six years. The IEA attributes about half of the net increase to accelerated (mostly AI) servers. Source, IEA, Energy and AI.
The IEA's projection puts global data-center electricity demand at roughly 415 TWh in 2024, rising to about 945 TWh by 2030. That is close to a doubling in six years, and the IEA attributes about half of the net increase specifically to accelerated servers, which is mostly AI. For a sense of scale, 945 TWh is on the order of a large industrial economy's entire annual electricity use, dedicated to running data centers.
This is where the software framing quietly breaks. You cannot scale a physical input the way you scale a download. Adding capacity means new substations, new power-purchase agreements, new siting fights, sometimes new generation that takes years to stand up. The queue to connect a large new load to the grid is measured in years in several major markets, not weeks. Power is not a line item you turn up. It is the ceiling the whole business runs into.
And it doesn't scale on the timeline software people are used to. A data center can be built in a year or two; the transmission to feed it, and the generation behind that, run on the timeline of heavy infrastructure, which is closer to a decade. So the gating resource isn't the building or even the chips inside it. It's the interconnect: the point where a gigawatt of new demand meets a grid that was planned for a slower century. That mismatch is why you now see AI companies signing directly for their own power, reopening plants, and negotiating for generation the way an aluminum smelter would. Those are not the moves of a software business. They are the moves of heavy industry, because at this scale that is what it is.
Which means the thing capital is actually buying, underneath the valuation, is bounded by megawatts. Not by how clever the model is. By how much power a company can secure, and how much useful output it can wring from each megawatt once it has it. The competitive question stops being "who has the best model this quarter" and becomes "who can turn a secured gigawatt into the most useful work," which is a throughput question, measured per megawatt, not per dollar.
Two ways to read what a company "has"
Question
Software framing
Production framing
What's the scarce resource?
Talent and model quality
Secured power, then throughput per megawatt
How do you grow output?
Ship code, add users
Add megawatts, or get more work per megawatt
What caps the business?
Market size
The grid, siting, and power contracts
What's marginal cost?
Roughly zero
A metered quantity of energy
That right-hand column reframes the competitive game. If output is bounded by power, then the axis that matters is not scale for its own sake. It is efficiency: how much useful work you get per megawatt. And that turns out to be something the hardware industry already measures, out loud, on its roadmaps.
If electricity is the input and tokens are the output, then the natural unit of this economy is work per watt. I've taken to calling it the kilowatt-token, and I want to be clear that this is a plain descriptor, not a coined bit of jargon and definitely not something we're trying to trademark. It is a name for a quantity the chip vendors already optimize on purpose.
Look at how NVIDIA talks about its own hardware. The metric on the slides is tokens per second per megawatt: how many tokens of inference the fleet can produce for a fixed power budget. By their benchmark, that figure improves on the order of a million times across six GPU generations, from Pascal in 2016 to Rubin in 2026.
Relative inference throughput per megawatt, about a million times over six GPU generations (Pascal 2016 to Rubin 2026). NVIDIA benchmark, workload-dependent, and shown as a conceptual redraw for editorial use. A vendor efficiency figure is not a universal measurement. Source, NVIDIA Technical Blog.
Two honest caveats, because a vendor number deserves them. This is NVIDIA's own benchmark, and it is workload-dependent: the figure you actually hit depends on your model, batch shape, and serving setup, and it will not be a clean million for most real workloads. And the chart is a conceptual redraw of the shape they publish, not a measurement I ran. So take the exact value as vendor-reported and the direction as the point: the industry is pouring its hardest engineering into tokens per second per megawatt, because that is the number that decides how much cognitive output a fixed power budget can buy.
That's the tell. When the most capital-intensive company in the supply chain organizes its roadmap around work per watt, "kilowatt-token" isn't a metaphor a copywriter invented. It's a description of what is already being built.
And notice the strategic consequence. If the constraint is power and the lever is efficiency, then the competitive axis is not who can raise the most money to buy the most chips. It's who gets the most useful work out of each megawatt. Scale still matters, but scale without efficiency just hits the power ceiling sooner. Efficiency is what moves the ceiling.
Play it out. Two companies, same secured power, say a gigawatt each. One squeezes twice the useful tokens per megawatt out of its stack: better silicon, better serving, less waste. It doesn't have a slightly better product. It has twice the factory on the same power contract. In a world where power is the thing you can't just buy more of on demand, a 2x efficiency edge is worth more than a 2x funding edge, because the funding edge still has to queue for megawatts and the efficiency edge doesn't. That inverts the intuition from the software era, where the best-capitalized player usually just outspent the constraint. Here the constraint is physical, and you can't outspend a grid interconnection queue. You can only out-engineer it.
This is also why the model itself is not where the durable advantage sits, or not only there. Model quality is real and it moves, but it's the most copyable layer, and everyone is drinking from the same few wells of talent and data. The efficiency layer, tokens of useful work per watt, compounds across every request forever and is far harder to replicate, because it's hardware plus systems plus operations, not a single artifact you can distill. If I were pricing one of these companies, I'd spend less time on this quarter's benchmark scores and more on the slope of their work-per-watt.
Here is the payoff, and it's where the raw numbers stop mattering and the right number starts.
Tokens per second per megawatt is the hardware's version of the metric. But tokens are not the product. A model can produce an ocean of tokens that resolve nothing: wrong answers, abandoned plans, confident work that gets thrown away. Raw token volume is a vanity number, the AI equivalent of counting lines of code. The quantity that actually matters one layer up is useful cognitive work per kilowatt-hour.
The production chain behind AI economics. The terminal metric is not raw token volume but useful cognitive work, and ultimately displaced human effort, per kilowatt-hour. The studio's stance stays augment, never replace; displacement here is the macro variable being priced, not a pitch.
Walk it up one rung at a time. Tokens per kWh is a hardware fact. Useful work per kWh discounts that for quality: how much of the token output actually turns into a merged change, a correct extraction, a decision someone could act on. And at the macro scale, useful work per kWh resolves into human effort displaced per kWh, the variable the economy is pricing. (Firewall still up: that last rung is an economic observation, not the thing I'd sell you.)
Now re-read the two-trillion valuation through this lens, and it stops being mysterious. The market is not paying a software multiple. Whether it knows it or not, it is pricing an expected future rate of cognitive work per kWh, integrated over the power these companies can secure. A bet that a firm will convert energy into useful cognitive output more efficiently than the alternatives, at enormous scale, for a long time. That is a physical-industrial bet dressed in a software cap table. It explains the velocity, too: a bet on a conversion rate can re-rate far faster than a business built on slowly compounding near-zero marginal cost, because the thing being priced is a slope, not a stock.
It also tells you which valuations are actually fragile, and it turns the vibe of the bubble debate into something you can actually check. Not the ones with big absolute numbers. The ones whose implied cognitive-work-per-kWh can't be reached with any plausible efficiency curve and any securable amount of power.
You can interrogate that, at least roughly. Take a valuation. Back out the annual economic output it implies at some sane multiple. Divide by the useful cognitive work one unit of that output represents, and you get the volume of useful work the company has to be producing at maturity. Now divide by the power it can plausibly secure over that horizon, and you get the work-per-kWh it has to hit. Compare that number to where the efficiency curve is actually heading. If the required work-per-kWh sits inside the plausible envelope, the valuation is aggressive but coherent. If it requires an efficiency the physics and the roadmap don't support, no story about the market can save it. The number is doing the arguing, not the mood.
I'm not going to run that calculation on a named company here, because doing it honestly needs inputs I'd have to source and caveat properly, and half-doing it would be exactly the kind of confident hand-wave this essay is against. The point is that the frame makes the question answerable in principle. "Is it a bubble" is a question about sentiment. "Can this valuation's implied work-per-kWh be reached" is a question about a production system. Only one of those has an answer you can check.
One shortcut is worth refusing on the way out, because it quietly reinstalls the wrong model: "AI is the new Internet."
It undershoots, and you can see why once you put the ages side by side.
Why the Internet analogy undershoots
Age
Input
Output
Industrial
Energy
Mechanical work
Internet
Information
Distribution
AI
Energy + information
Cognitive work
The industrial age converted energy into mechanical work, and its constraints were physical: fuel, machines, throughput. The Internet age converted information into distribution, and its constraints were the ones we all learned to reason about: near-zero marginal cost, network effects, winner-take-most. When people say "AI is the new Internet," they are importing the second row's economics, the free-copy, zero-marginal-cost intuition, onto something that doesn't obey it.
AI is the row that folds the other two together. It takes energy and information as inputs and produces cognitive work as output. It has the physical constraints of the industrial age, power, siting, throughput, and the information dynamics of the Internet age at the same time. That is why the pure-software analogy keeps misfiring. The correct reference class isn't the web. It's the web running inside a power plant's constraints.
I genuinely don't know, and I've stopped thinking it's the interesting question. Valuations correct or they don't; specific companies are mispriced in both directions right now; some of today's names will look obvious in ten years and some will be case studies. All of that can be true and none of it touches the thing underneath.
The question that survives the cycle is physical, and it's the same in a boom and a bust: how efficiently do capital and electricity convert into useful cognitive output? Bubbles pop. The efficiency question stays, because it's a property of the production system, not of the mood in the market. Whoever wins the next decade of this wins on cognitive work per kilowatt-hour, on getting more real output from each megawatt, not on who told the best story about it.
That reframe is not academic for me. It's the altitude I actually work at.
I build and operate software with agent fleets, and my whole job, once the model is chosen, is the conversion efficiency in the small: turning tokens into work that's actually correct instead of tokens that merely look busy. Grounding agents in canonical context so they don't burn compute rediscovering what's already known. Gating the irreversible actions so a confident mistake doesn't become an incident. Instrumenting the runs so I can see where the useful work is and where the waste is. Measuring the output, not the activity. That is the kilowatt-token question scaled down to one team's systems: how much useful work comes out per unit of compute spent, and how do you raise it.
The macro version and the micro version rhyme. The economy is pricing cognitive work per kWh at the scale of continents. On a team, you're doing the same accounting on the scale of a sprint. Same question, same firewall: measure the work that comes out, and keep augment and displace in separate columns.
None of that is a metric we sell you. Work per kWh is the lens I read this whole shift through, and the studio's job is the unglamorous version of it: we build and operate software with agent fleets, and we measure the useful work that actually comes out.
Bubble or not, that's the question that will still be worth asking when the current one has been settled and forgotten.
We help engineering teams get more useful work out of their agents for less wasted compute, across reliability, observability, context, and cost. Measured, not promised.
Bubble-or-not is the wrong frame. Both the bubble camp and the believers treat AI as pure software economics: multiples, hype cycles, a marginal cost near zero. But AI runs on a physical chain of capital, electricity, and compute, and a token is not free to produce. The question that survives the cycle isn't whether valuations correct. It's how efficiently capital and electricity convert into useful cognitive output. Bubbles pop; the efficiency question stays.
Why does a two-trillion-dollar valuation happen so fast now?
By one investor scenario, a leading AI company could be valued near $2T roughly 5.8 years after founding, against 23 years for Alphabet and more than 44 for Apple and Microsoft. That specific number is a projected scenario for a potential 2026 IPO, not an achieved public-market valuation. What the market is pricing is not a software multiple but the expected future rate at which these systems convert energy and capital into economic work.
How much electricity will AI data centers use?
The IEA projects global data-center electricity demand rising from about 415 TWh in 2024 to roughly 945 TWh by 2030, close to a doubling in six years, with accelerated (mostly AI) servers accounting for about half of the net increase. That makes electricity a binding constraint on AI output. What capital is really buying is bounded by megawatts: grid capacity, siting, and power contracts, not just GPUs.
What is the 'kilowatt-token'?
It's a plain descriptor for the unit the AI economy actually runs on: work per watt. If electricity is the input and tokens are the output, the meaningful efficiency measure is tokens per second per megawatt. NVIDIA reports its own benchmark improving on that order by roughly a million times across six GPU generations (workload-dependent). The term names something the industry already engineers for; it is not a coined metaphor or a trademark.
Does 'human effort displaced' mean AI replaces developers?
No. In this essay 'human effort displaced' is a macro-economic observation, the variable the market is pricing, not a product pitch. The studio's stance is augment, never replace: we build and operate software with agent fleets and measure the useful work that comes out. Displacement is what the economy is pricing at the aggregate; it is not a recommendation to remove your engineers.