Where Does The Money Actually Go? A Recap of Stanford's MS&E 435

Jensen Huang's five-layer cake goes up on the projector twice in MS&E 435. Energy, chips, infrastructure, models, applications, with company logos stacked on each layer: Constellation and GE Vernova at the bottom, NVIDIA and Broadcom above them, then AWS and CoreWeave and Crusoe, then OpenAI and Anthropic, then Cursor and Harvey and Waymo at the top. Two of those logos belong to companies whose founders teach a week of this course. Both times the slide appears, the guest gets a version of the same question: you have $100, which layer do you put it on?

Ali Ghodsi (Co-Founder and CEO of Databricks) runs a company that sits on the infrastructure layer of that slide. His answer was that he's a computer scientist, not an investor. Apoorv Agrawal (Partner at Altimeter Capital), who teaches the seminar, pressed him anyway. He picked early-stage applications.

That exchange is the course in miniature: nine practitioners, one slide, and no agreement about where the profit ends up.

Stanford published all 9 lectures from MS&E 435: Economics of the AI Supercycle, Agrawal's seminar, which brings in one practitioner per week from a different layer of the AI stack. This post walks the lectures in course order: the economics of generative AI, the GPU economy, AI factories at gigawatt scale, enterprise software, a frontier lab's compute strategy, enterprise internal knowledge, inference as COGS, coding agents, and AI in life sciences.

One note on that order, because it isn't mine. The syllabus is built bottom-up: week 2 is silicon, week 3 is data centers, weeks 4 through 6 are enterprise software and models, weeks 7 through 9 are applications. The lecture order is the stack.

Where is the money in AI today?

The first lecture puts up two versions of the same three-layer pyramid. In cloud software, applications earn about $600B a year, infrastructure about $300B, and semiconductors about $80B. In AI, it's $60B, $75B, and $300B. Same three layers, upside down. The gross profit split is the sharper version of the same picture: semiconductors took 87% of AI gross profit in 2024 and 79% in 2026, while applications moved from 3% to 7%. In cloud software, applications take 70%.

Agrawal's explanation is marginal cost, and every later lecture runs into it. Traditional software sold the next seat for almost nothing, which is how those businesses ran at 80–90% gross margins. Every AI request spends GPU time instead, and somebody has to own the GPU. So the question the seminar puts to nine practitioners over nine weeks is whether that inversion is a phase or a structure, and each of them answers it from the layer they happen to own.

1 - Where The Money Is, And Isn't

Between his 2024 analysis and his 2026 update, NVIDIA added roughly $175B of revenue. The entire application layer of the AI economy (OpenAI, Anthropic, and every other company selling an AI product) is about $60B. One vendor's two-year growth is close to three times the size of the layer everyone expects to win.

Agrawal has been publishing this analysis since before the course existed, which is what makes lecture 1 more than a syllabus walk. The Economics of Generative AI came out in April 2024 and priced the whole stack at about $90B: $75B in semis, $10B in infrastructure, $5B in applications. He called that shape A-shaped, against the V-shaped distribution of the cloud stack, and put the industry in what he labeled Inning #1. The expectation was migration upward. Two Years Later is the scorecard, and he grades himself honestly: the ecosystem grew 5x to roughly $435B, and the shape didn't move.

The margin table underneath is the part I'd hand to anyone who only has ten minutes with this material. In 2026, semis earned about $225B of gross profit at 73% margins, infrastructure about $40B at 55%, applications about $20B at 33%. Two years earlier NVIDIA was running at 85% and the app layer at 50–55%. So chip margins came down by roughly a dozen points, and the layer still kept 79% of the profit pool. The app layer grew 12x and lost close to 20 points of margin doing it. The apps are buying their inputs from the layer they're supposed to displace.

Which sets up the number that reframes the whole debate: the semis share of gross profit moved 8 points in two years. His own read is that app-layer dominance would take more than a decade at that rate. And his one-line summary of the 2026 picture is the sharpest thing either essay says: "Semi is a one-player game. Apps is a two-player game. Infra is the only competitive layer."

The lecture then lists what would actually move value up the stack, and it's four things, not a sentiment.

  1. Custom ASICs that work in production.

  2. A change in hyperscaler capex guidance.

  3. Real price competition in inference.

  4. Vertical integration, the way Google, Apple and Meta captured layers in earlier cycles.

There's a fifth question he leaves as a question: whether AI startups become platforms or end up as features inside AWS.

The last stretch turns to consumers, and the tiers are worth memorizing because they set the ceiling on any consumer AI business. Core utilities like Chrome and WhatsApp live at 2–3 billion weekly actives. Social platforms (Facebook, Instagram, TikTok) sit at 1–1.5 billion. Niche apps run 300–600 million, which is where Spotify is. ChatGPT is around 900 million weekly actives, Gemini 200–250 million, and in the lecture Agrawal's phrasing is that ChatGPT "has just overtaken the niche category." The engagement data in his consumer AI series is the more interesting half: 66% week-4 retention against Gemini's 44% and Claude's 41%, and a 45% DAU/MAU ratio against Gemini's 22%. Habits, not headlines.

Then the monetization fork, where he's willing to make a call the lecture only gestures at. About 95% of ChatGPT's users don't pay. At Google's global $84 of annual revenue per user and Meta's $57, his estimate of $30 per free user per year works out to roughly $25B, which would make advertising the larger opportunity for consumer AI than subscriptions. Google can keep Gemini ad-free because a $295B search business pays for it. OpenAI has no such cushion.

What lecture 1 doesn't do is finish its own sentence. It names the four things that would flip the stack and declines to bet on any of them. That restraint is the design of the course: the same question goes to eight more people, each of whom owns a different layer of the answer.

2 - The GPU Economy

In September 2025, Jonathan Ross went on 20VC to argue that OpenAI and Anthropic would build their own chips, and to field a quickfire question about whether NVIDIA reaches $10 trillion. Ross had led the team at Google that built the first TPU before founding Groq in 2016. Three months after that interview, NVIDIA paid about $20B to license Groq's inference technology and hired Ross along with Groq's president, Sunny Madra. By April 2026 Madra was sitting in a Stanford classroom as an NVIDIA employee, describing how fast it happened: he showed Jensen Huang a working system, and by his account the deal took "probably just over a month."

Brad Gerstner takes the first twenty minutes before any of that comes up, and the slides are an investor's argument rather than an engineer's. Technology has gone from 5% to 15% of US GDP. Tech returned 15% a year over fifteen years against 6% for everything else. Automating 10% of knowledge work is worth $5T a year. Gerstner has been making this case for a long time. Altimeter grew from a $3M Boston hedge fund into a $15B asset manager on a concentrated portfolio of technology companies, public and private alike. The claim underneath all three slides is the same: the share of the economy running on software keeps going up.

Then Madra takes the chip question, and his answer is the most useful thing in the hour for anyone evaluating a hardware startup. Groq's LPU is a genuinely different design: deterministic, single-core, memory units interleaved with compute, no caches and no branch predictors, with execution controlled entirely by the compiler. Cerebras built something equally unusual. Both companies had a hard time anyway, because the thing being sold isn't a chip. It's racks, interconnect, a software stack developers already know, and enough supply to promise someone a cluster next quarter. A fast chip is not a business.

The economics of the hour run on tokens. Reasoning models and agents spend orders of magnitude more tokens per task than a chatbot answering once, which is why falling prices don't reduce the bill. Gerstner puts the decline in inference cost at roughly 90% over the past year and something closer to 99% over two and a half, and the sources are unglamorous: TSMC's supply chain, advanced packaging, quantization, and architectural work like disaggregating prefill from decode. Here's the part that makes the "cheaper means less spending" intuition collapse. H100 spot prices have been going up through all of that deflation, because demand keeps outrunning whatever the supply chain delivers. Cost per token falls; total spend rises.

The bubble question gets answered with one number rather than an argument. Anthropic added something like $10B of annualized revenue inside a single month in March 2026, which Gerstner sets next to the combined annual revenue of Databricks and Palantir. He's an investor in the company and these aren't audited figures, so treat them as a practitioner's report and not a filing. The point he's making with it is about sales motion, not size: that revenue didn't require a million salespeople convincing a million companies one at a time.

What the last stretch does is stranger, and it's the part I keep thinking about. Gerstner says the people building these systems (he names Dario Amodei, Sam Altman and Elon Musk) have told him privately that AGI has either arrived or is close. Nobody in the room defines AGI. From there the conversation turns to who receives the wealth if any of that is true, which is where his Invest America work comes up. Then to what work looks like when problem-solving gets cheap and persuasion, leadership and taste are the scarce parts. Three weeks later a different guest will put up a slide whose stated goal is to make the human the bottleneck. Both of those are in the same course, and nothing in it joins them.

3 - Building AI Factories

A rack in a legacy cloud data center draws 2–4 kilowatts. A current AI rack draws 140, and Chase Lochmiller (Co-Founder and CEO of Crusoe) expects the next generation somewhere between 600 kilowatts and a megawatt. He gave those numbers on a McKinsey podcast in November 2025, and they explain why "data center" is the wrong word for what week 3 is about. The building is a power plant with a computer bolted to it.

Agrawal sets the scale first. The Big 5 (Oracle, Meta, Google, Microsoft and Amazon) are projected to spend $645B on capex in 2026. Then Lochmiller's slide puts up an equation. Data + Algorithms + Compute + Energy + Data Centers = AI. Crusoe works the last three, and his framing of the output is electrons in, tokens out.

The scarce input is a powered data center, not a data center. Grid interconnection queues run years, as Eric Flaningam lays out in A Primer on AI Data Centers, so Crusoe inverted the problem: stop waiting for power to reach the compute and put the compute where the power already is. The company started by capturing natural gas that was being flared at oil wells, which matters because methane traps about 84 times more heat in the atmosphere than CO2. Lochmiller phrases the founding question this way: why don't we bring a market to the gas? West Texas wind came next, where curtailment runs 25–30% and power sometimes prices negative.

Abilene is the worked example, and it's Stargate's first site. The number that makes the strategy concrete there isn't a cost. Crusoe delivered the first 100 megawatts there in 11 months against a 12-month target, while 32 other bidders on the same project proposed two and a half years or more. The campus is headed to 1.2 gigawatts, powered by a 350-megawatt gas plant alongside the wind. At peak it ran 5,800 workers a day in a city of 120,000, about half of them hired locally. Lochmiller mentions another campus where the on-site staff will outnumber the town's residents.

Then the lecture turns into a spreadsheet, and this is the part I'd keep. The unit is dollars per megawatt. The building costs about $19.2M/MW, and the largest single line is labor at $4.7M, ahead of materials at $3.9M and the on-site gas plant at $3.0M. The IT inside adds about $40M/MW, of which $30M is GPUs. Three quarters of the IT bill is one line item. Call it $59M/MW up front and roughly $1.1M/MW a year to operate. Rented out as raw infrastructure, a megawatt returns about $15M a year; sold as managed AI services on the same hardware, up to $30M, which the slide calls a 100% revenue premium. Run the arithmetic and the infrastructure case pays back a little over four years, the managed-services case just over two. Crusoe has 3.5+ GW under construction or contract, including 1,000 MW across six buildings in Claude, Texas. It also sells a modular product built out of what he calls data center Lego blocks: Crusoe Spark, 500 kW air-cooled units, delivered in as little as three months, 400+ already deployed.

Here's where I'd push back on the course's own consensus. Three guests at three different layers tell you that compute scarcity is effectively permanent: Lochmiller with gigawatts under construction, Sachin Katti in week 5 with OpenAI's 30 GW target against a national build-out near 100 GW, and Tuhin Srivastava in week 7 with an airport security line that never gets a quiet hour. Utilization data is on their side and I believe the demand claim.

But scarcity is doing work that an asset-life estimate should be doing. The entire case for $59M/MW rests on those GPUs staying economically useful long enough to earn it back, and the one person asked directly about the depreciation curve says forecasting is hard and moves on. That question came from a student in this same lecture. His answer concedes the mechanism: older compute commoditizes while cutting-edge systems keep a premium. That's a depreciation schedule described and not quantified. Worse, the revenue figure carrying the payback math has a footnote, and the footnote is a SemiAnalysis estimate of an April 2026 five-year GB300 contract price. Five years of assumed contract against four years of payback is the whole margin of safety. That's thin.

So my position: "compute is scarce" has become the industry's substitute for a depreciation estimate. It's a demand statement standing in for an asset-life statement, and the two fail in different ways. If demand keeps growing and the hardware holds its rental value, $59M/MW is cheap. If demand keeps growing and each GPU generation halves what the last one rents for, the scarcity sentence stays true and the investment still doesn't work. Nine lectures, and the number that decides which of those worlds we're in never appears on a slide.

The strongest version of the other side is Crusoe's own, not the industry's. An operator that owns the gas plant, the land and the substation holds a cost advantage that survives a hardware cycle, because the wing full of GPUs is the part that depreciates and the power is the part that doesn't. Racking new accelerators into a shell you already own is a different investment from building the shell. That's a real answer, and it's why I'd rather own Crusoe's position than a pure GPU rental business. It still isn't an answer to what a megawatt earns in year six.

4 - Infrastructure, Enterprise

"Majoring in computer science today will be like majoring in journalism in the late 90's." That line is from Chris Paik (Co-Founder and General Partner at Pace Capital), in a June 2024 essay called The End of Software that sits on this week's reading list. The argument takes three sentences. Software was expensive because developers were expensive, and developers were expensive because they were skilled translators between human language and computer language. LLMs do that translation, so the cost of creating software goes to zero. What follows is a Cambrian explosion of disposable software in which "Salesforce will not be replaced by another monolithic CRM" but by "a constellation of things that dynamically serve the same intent and pain points."

The market half-believes it. Agrawal's chart puts software's median EV/revenue multiple at 3.3x in 2026, the low point on a slide titled "Revenue Multiples All-Time Lows," down from 6.4x in 2018, with the three drops annotated: Taper Tantrum, Inflation and Rates, AI Headwinds.

Ali Ghodsi's first move is a reductio. If all software is dead, he asks, then isn't OpenAI dead? Isn't NVIDIA? He throws in SpaceX. Every company anyone names as the future of anything is a software company, which is a cheap shot and also correct.

His real claim is narrower and more useful. Three things fall at the same time: the cost of building software, the barrier to entry, and the switching cost of leaving a vendor. Competition gets brutal, which is the mechanism behind the app-layer margin compression in Agrawal's own numbers. What survives isn't the software. It's whatever moat was never software to begin with, and he tells the room to go read Hamilton Helmer's 7 Powers before naming the one he's most confident in: economies of scale.

Then the part he's since built a whole keynote around. Two months after this lecture, opening Data + AI Summit in June 2026, his line was that AI doesn't have an intelligence problem, it has a context problem. The classroom version uses his own support organization as the example. Resolving a support case requires knowing why the product behaves the way it does, which exception was granted to which customer, and who to ask: the kind of thing an employee accumulates over years and an agent has no access to. Databricks has roughly 20,000 customers. That's a lot of history nobody wrote down.

The best story in the hour is about a connector. Databricks builds integrations to pull data out of other systems, the process was slow, and Ghodsi wrote one himself in two days to make a point about what AI could do. Then the organization's reaction taught him the actual lesson: the two days weren't the bottleneck. The process for producing connectors was, and fixing it meant rebuilding that process from scratch, changing who was on the team and sending some of the work outside. The demo proved AI could write the code. It took a reorganization to make that matter.

He also declines, gracefully, to play the game the course keeps setting up. Handed the five-layer cake and $100, he says he's a computer scientist rather than an investor, gets pressed, and picks early-stage applications, from the seat of a man running an infrastructure company at a $100B+ valuation. Asked what he uses himself, he names Cursor and Databricks' own Genie, and says the open-weight models are close enough now that Moonshot's Kimi is a real option.

One thing I can't resolve. The session is titled "Service as a Software," a phrase that implies the entire selling motion of enterprise software is about to change, and neither he nor Agrawal ever defines it. 39 minutes on whether software survives, and the most aggressive claim in the room was on the title slide.

5 - Infrastructure, Capstone Case

OpenAI's compute capacity more than tripled in 2025, to roughly 1.9 gigawatts. Its revenue tripled too, and the chart Agrawal opens week 5 with lays those two lines on top of each other. His word for it is hypercorrelated. The guest asked to explain the chart is Sachin Katti, who spent six months as Intel's CTO and then joined OpenAI in November 2025 to run industrial compute under Greg Brockman. The course assigns a keynote he gave in the Intel job, Scaling AI at the Speed of Openness: From Silicon to Systems. He's now on the buy side of the same problem.

The targets are the part worth writing down. OpenAI's aspiration is 30 GW, split between research and products, so that researchers aren't rationed when they want to try something. Katti puts the planned US hyperscaler build-out near 100 GW. Sam Altman has floated a gigawatt of new infrastructure per week as a long-term goal, which is the kind of number that stops meaning anything, so hold onto the 1.9 you started with.

There's a detour into Intel that neither of them can quite play straight. The slide shows the stock up 370% from its May 2025 low, at a $481.6B market cap as of the day of the lecture. Katti's read: a global shortage of manufacturing capacity, plus CPUs mattering again inside AI systems. The interviewer jokes about the timing of his departure. He laughs and doesn't take the bait.

Custom silicon is where the course's threads start tying together. Two weeks after Jonathan Ross went on a podcast predicting that OpenAI and Anthropic would build their own chips, OpenAI and Broadcom announced 10 gigawatts of OpenAI-designed accelerators, with racks deploying from the second half of 2026 through the end of 2029. OpenAI designs the accelerators and systems; Broadcom builds them and wires them with Ethernet. Brockman's line in the announcement explains why a lab would bother: "By building our own chip, we can embed what we've learned from creating frontier models directly into the hardware."

The demand side gets a slide that doubles as OpenAI's roadmap. Chatbots in 2023, reasoners in 2024, agents in 2025, innovators in 2026, then organizations; with a note that GPT-5.3-Codex "was the first model instrumental in creating itself." Each step up that ladder makes the unit of work heavier. A chatbot answers once. An agent runs a long graph of calls, holds context across them, and waits on tools.

Which brings the hour to its most teachable moment. Press enter in a ChatGPT-style interface and about 500 milliseconds pass before the first token lands. Asked where that budget goes, Katti says most of it is prefill: the model has to page in the entire context, a whole codebase in the coding case, before it can emit anything. Then the lesson that generalizes. Put in fast-inference silicon like Cerebras and you don't get a 10x product; you expose whichever layer was second-slowest. Hence his framing that time to compute matters more than the amount of compute, and hence gigawatt campuses instead of 50 MW sites.

The third assigned piece is a Dylan Patel (Founder of SemiAnalysis) interview published six days before this session, and it's the strongest version of the case my depreciation skepticism has to beat. His framing is that nobody in this market loses margin to competition; they lose capacity to undersupply. Labs are selling out ahead of what they can serve, TSMC's 2026 capex guidance of roughly $57B could be on a path to $100B by 2028, and memory demand is structural rather than cyclical because inference runs hot around the clock. If that's right, the depreciation question I raised matters less, because everything gets used at full utilization until it dies.

The rapid-fire at the end is funnier than it is informative. First company to $10 trillion? "The easy answer is NVIDIA, right?"

One slide from this lecture has stayed with me, and not in a good way. Its title is the goal: make the human the bottleneck. The success condition underneath reads that the user should be waiting on their own decisions, not on model inference, tool hops, or scheduler delays. Three weeks earlier this course spent its closing minutes on what people will do for a living when problem-solving is cheap, and who receives the surplus. Now the design target is a person who never waits, in a session that never mentions the earlier one.

6 - Enterprise Internal Knowledge

DeepSeek's R1 spent about 150,000 H800 hours on its reinforcement learning run. V3's pre-training spent 2.4 million. By Yash Patil's numbers, the step that makes a model useful costs roughly 5% of the step that made it smart.

Patil went from Stanford to OpenAI, worked on post-training and reasoning models there, and left to found Applied Compute. His lecture hangs on one table, and it's the artifact I'd keep from the whole course: every era of AI, its bottleneck, and the thing that broke it. Before 2012, hand-designed features, until deep learning started learning features on its own. 2012 to 2016, data and GPUs, until ImageNet and AlexNet proved neural nets could scale. Before 2017, sequence architecture, until transformers. 2018 to 2021, compute and scale, until GPT-style models showed that bigger got broadly capable. 2022 to 2024, usability and alignment, until RLHF and ChatGPT made models usable by normal people. The current row doesn't say compute. It says high-quality tasks, broken by evals, verifiers and RL environments.

Andrej Karpathy's 2025 LLM Year in Review reaches the same place from the research side. His shift of the year is RLVR: training against rewards a machine can check rather than preferences a human has to rate. His framing of the application layer is the one that matters here: foundation models arrive as generally capable graduates, and the companies that turn them into working professionals do it with private data and domain-specific feedback. He also has the best available word for why this is uneven. "Jagged intelligence": polymath and confused beginner in the same system, which is the chart Agrawal put up in week 4 under the title AI's Jagged Frontier.

Code and math got there first for one reason: a machine can mark the homework. Almost nothing else in an enterprise works that way, and that gap is Patil's whole business. The scarce input isn't intelligence, it's a written-down definition of what good looks like, which is why he says evals set the training roadmap rather than measuring it afterward.

The worked example is DoorDash, which onboards more than 100,000 merchants a year. The slide says it trained a proprietary agent using internal experts and got a 30% relative reduction in critical menu errors. What made that possible wasn't in any model. It was in the heads of the people who know which modifiers and add-ons are allowed, and it stayed there until somebody turned it into training signal. Production models can then be small and fast: his example is Cognition/Windsurf checking a user's code for bugs in under two seconds.

Continual learning is where the lecture points next. The slide's own vocabulary is a mix of weight updates, context and harness, with a note that this capability is owned by companies themselves. Beside it sits a chart of Cursor's Composer improving through real-time RL. That's the enterprise translation of the other assigned reading. David Silver and Richard Sutton's Welcome to the Era of Experience argues that the era of human data is ending. High-quality human text in math, coding and science is running out, and imitation can't exceed what people already know. Their alternative is agents living in streams of experience over months or years, earning grounded rewards from outcomes rather than human ratings. Their example is AlphaProof, which started from around a hundred thousand formal proofs and then generated a hundred million more. In the DeepMind version the stream is a lifetime. In the enterprise version it's your telemetry, and the reward is whether the merchant's menu came out right.

Agrawal tries to open an architecture fight near the end, quoting Ghodsi's line that airplanes aren't as efficient as birds and fly anyway. Patil doesn't take it. Scaling transformers works right now, he says, and a different architecture starts to matter if scaling stops working.

What the table doesn't have is a column for who pays. The current bottleneck is evals, verifiers and RL environments, all of which somebody has to build, for every task, in every company, mostly from scratch. Patil's startup sells exactly that. Whether that's the answer to the question or the reason the question got framed this way, the lecture doesn't say, and I can't tell from the outside.

7 - Applications, Applied AI

Two people say 95% in this course and mean opposite things. Agrawal's framing in week 7 is that 95% of tokens run on frontier models and 5% on post-trained or open ones. Tuhin Srivastava, who co-founded Baseten in 2019 and runs it, has said on No Priors that 95% of the tokens Baseten serves come from models its customers have modified. Both numbers are probably right. One is the market; the other is the subset of the market that shows up at a specialized inference cloud, which is a selected sample and shouldn't be read as a forecast. Worth knowing before the lecture's central claim arrives, because that claim is a vendor's.

The claim itself is the best sentence in the session: inference is the COGS of AI value delivery. Serving a model isn't a fixed platform cost, it's the cost of goods sold against every request, so an AI product's gross margin is set by what its tokens cost and how fast they arrive. Baseten's own explainer on inference is useful on why those two pull against each other. Time to first token is the number users feel; raising concurrency lifts your throughput per GPU and hurts that number. Cheap tokens are never free. You pay in latency or in engineering.

The customers make it concrete. Wispr Flow runs custom speech-to-text models. Abridge, whose ambient scribes sit inside hospital EMRs, runs about 20 different models in production. What makes the Abridge example more than a logo is what Srivastava says elsewhere about why that business is defensible: every clinician edit and every downstream action in the record is a signal no frontier lab can see. Private feedback, not a private model.

Then the slide the section turns on. Post-trained open-weight models, it says, increasingly deliver frontier-level performance at 30% of the cost, with open models trailing the frontier by under 90 days, and a post-trained Kimi giving Sonnet-level intelligence at Flash-level costs. Sitting in the room, the thing to notice is whose slide it is. Baseten sells post-training and the inference behind it, and there is no eval on the screen. Take the direction and distrust the decimal.

Even taking the number at face value, the arithmetic has an edge the slide doesn't show. A 70% cut in cost per token is worth 70% of whatever share inference holds in your COGS, and against that saving sits the fixed cost of building the custom model, writing the evals and maintaining both as the frontier moves. The saving scales with volume. The fixed cost doesn't. So post-training pays above some volume threshold and destroys value below it, which is what Srivastava means by his own rule: no post-training before product-market fit. Neither the lecture nor the interview says where the threshold sits. That's the number I most wanted from this session and didn't get.

The harder problem is one the framing hides. Acquiring 1,000 B200s, he has said, now means a three-to-five year contract with 20–30% of the total value prepaid. Baseten runs something like 90 clusters across 18 cloud providers at mid-90s utilization, which is the operational answer to that contracting reality: if you've committed to the capacity, the only thing left is to keep it full. And once capacity is prepaid on a multi-year term, inference stops behaving like COGS at the margin. You already bought it. For a startup serving its first thousand users, inference is a variable cost. For anyone large enough to be interesting, the bottom of the load is fixed cost with a five-year tail; the same unpriced asset-life question from week 3, wearing different clothes.

Srivastava is direct about what could break his company: open source, the application layer, and access to compute. On the third, the numbers are large. He expects Baseten's API business to need roughly 150,000 B200 equivalents within two years, around $7B of compute, and he puts owning infrastructure at about 30% cheaper than renting it. Baseten raised a $300M Series E at a $5B valuation in January 2026, with NVIDIA and Altimeter among the investors. Worth stating plainly: the seminar's instructor is a partner at one of the firms backing the company whose CEO is teaching the session.

Asked whether the compute crunch is a 12-month problem or a multi-year one, he compares it to the security line at JFK, which never gets a quiet hour because demand keeps arriving. His version of the punchline is that it's always 5 PM somewhere in the world.

8 - Applications, Coding AI

Guillermo Rauch (Founder and CEO of Vercel, and the creator of Next.js) puts up a market slide made of three rings. The inner ring is JavaScript developers. The next one out is business users who want software without writing it. The outer ring is autonomous agents, and that's the ring the company is building for now. He sums up where it leads as "the AWS of AI or agents." The footprint he starts from: Next.js at 33M+ weekly downloads, the AI SDK at 12M+, 17M+ people building on Vercel every day, 5M+ deployments a day.

The names for this changed three times in nine years. On February 2, 2025, Andrej Karpathy posted the tweet that named vibe coding: "I 'Accept All' always, I don't read the diffs anymore... it's not really coding - I just see stuff, say stuff, run stuff, and copy paste stuff, and it mostly works." A year later he called that a "shower of thoughts throwaway tweet" and proposed agentic engineering instead: you are "not writing the code directly 99% of the time, you are orchestrating agents who do and acting as oversight." Back in 2017, his Software 2.0 essay ended by listing what the new kind of code still lacked: no IDE, no GitHub for datasets, no package manager. The naming keeps changing. What it points at doesn't: something has to catch whatever the machine produces.

That's Rauch's pitch in one line. Writing code isn't shipping it. Getting generated code into a running, secure, observable service is the part that stays hard, and his claim is that it needs a different kind of cloud. The "Traditional Cloud vs. Agent Cloud" slide is the best artifact in the lecture, and it's a table of swaps. Static pages become generative, streaming interfaces. The CDN delivers tokens instead of pixels. Compute stops running quick, human-written work in the foreground and starts running long, IO-bound, agent-written work in the background. Vercel's description of itself runs on the same three beats: infrastructure for coding agents, to ship agents, automated by agents.

His business conclusion is that tokens become the commodity and pricing moves from SaaS seats toward usage. Run that from the other end and it gets more interesting. An agent produces more deployments per hour of human attention than a person typing does, so the count of deployments grows faster than the count of developers. And a deployment isn't a token. Serving a cached response costs bandwidth and a cold start. Serving a token costs GPU time. That difference is week 1's marginal-cost story, and it's what put classic software at 80–90% gross margins and AI applications at 7% of the stack's gross profit. The deployment layer never took that hit. So the coding-agent boom can put a business with software's old margin structure directly on top of one paying for silicon by the request. Rauch says tokens are the commodity. That is the same sentence read from the side of the trade that benefits.

The evidence on his slides is all about defaults. One says Vercel is Claude Code's default deployment target, picked in 86 of 86 frontend deployment decisions. Another shows deployments tripling between October 2025 and May 2026. A Business Insider excerpt describes Meta Superintelligence Labs steering staff toward platforms like Vercel. Read 86 of 86 as distribution won and it's the strongest slide in the hour. Read it with one property of the new customer in mind and it turns over. On Sequoia's Training Data podcast, Rauch stated that property himself: "Your customer is no longer the developer. Your customer is the agent that the developer or non-developer is wielding." An agent has no habits, no procurement cycle, and no half-finished migration it would rather not redo. Developer lock-in was made of those three things. The property that made Vercel the default is the property that makes the default cheap to lose, and the default lives in someone else's product.

The tripling chart is the only direct evidence anywhere in the course that agents are already moving infrastructure demand. Nothing in the lecture separates deploys an agent triggered from the ordinary growth of a company that was already growing before agents existed. In the Q&A, Rauch draws the line he'd defend: off-the-shelf applications may get rebuilt, while the systems of record underneath them keep their value. Asked where value accrues across chips, models, infrastructure and coding agents, he talks about sandboxes, deployment and the token economy. He doesn't pick.

9 - Applications, AI in Life Sciences

The last session is the only one with two guests, and it carries the most concrete sentence in the course. It's on Chai Discovery's goal slide: "Generate antibody drug candidates ready for IND-enabling studies, entirely on the computer." Under it are three challenges. Bind drug targets with high affinity without library-based screens, meet developability properties, and reach therapeutic applications such as specificity to cancer mutations. Joshua Meier worked on ESM-1 at Meta and spent time at OpenAI before co-founding Chai. He describes the company as a computer-aided design suite for molecules, sold into the pharmaceutical ecosystem rather than used to turn Chai into a pharma company.

Eric Kauderer-Abrams of Anthropic supplies the arithmetic that makes the attempt worth making. A drug takes 10 to 15 years from idea to market, and the fastest on record is about 5 to 6. Roughly 30 net-new targets a year get pursued in the clinic. Against that sit maybe 10,000 potential targets and 19,000 genes in the human genome. Two orders of magnitude separate what biology offers from what anyone attempts.

Neither of them claims a model closes that gap by itself, and the published numbers show why. Chai's Chai-2 preprint reports a 16% hit rate in fully de novo antibody design, from 20 or fewer designs per target across 52 targets. None of those targets had a preexisting antibody or nanobody binder in the Protein Data Bank. That's about a 100-fold improvement on previous computational methods, and it produced at least one hit for half of the targets. Read the denominator, though. The result is a single round of experimental testing, and every hit in it was called by a wet lab. On No Priors, co-founders Jack Dent and Joshua Meier put the loop at about two weeks per validation cycle. That floor is biological, not computational.

Which is where I'd push back on the framing this course keeps reaching for. The story told about AI in drug discovery is a speed story: compress the timeline, get medicines out faster. Their own number argues against it. The 5-to-6-year record was set without any of this, so a perfect front end lands you near a mark the industry already hit. And the back half of development is what both speakers describe as operations: patient recruitment, trial sites, manufacturing, regulation. None of that is a design problem. What 20 designs and one plate change is the cost of an attempt. Once an attempt is cheap, the limit on how many drugs get made moves from design capacity to trial capacity and capital. The bet here isn't on the calendar. It's on the count.

The honest version of the counterargument is that better candidates might fail less often once they reach the clinic, which would move the timeline by raising trial success rates rather than shortening trials. One round of testing on 52 targets says nothing about attrition three years later. That number doesn't exist yet for anyone.

What both guests keep returning to is the loop between a model and a physical experiment: AI wired into lab instruments, programmable chemistry, a "lab in a box." Kauderer-Abrams is skeptical of businesses built on selling tools to pharma, since falling barriers may let small teams run their own drug programs, and he'd rather own hard experimental capacity. He names Plasmidsaurus, Twist and Adaptive as the kind of AI-native wet-lab companies he likes. Worth holding that next to what his employer had already shipped seven months before the session. Claude for Life Sciences launched on October 20, 2025, with connectors into Benchling, PubMed, 10x Genomics and Wiley, and a Protocol QA score of 0.83 against a 0.79 human baseline. A general model with connectors isn't a discovery platform, so his position and the product can both stand. It's still the tool business he says he wouldn't bet on.

The session closes on scaling, and on a question neither speaker answers. Training runs are entering the billion-dollar range, with ten or a hundred billion imaginable. Nobody in the room is sure whether a model trained to fill in a masked paragraph understands biology or only predicts it well.

The tenth lecture

Nine practitioners got the $100 question, and nine of them answered from the layer they own. That's not a complaint: it's the only honest way to answer it, and it's why the guest list is worth more than any single hour of tape. But the number that decides the question never went up on a slide. Not once in nine weeks.

Crusoe's own figures pay back a megawatt in a bit over four years on raw rental, and in about two with managed services running on the same hardware. Those two numbers sit on opposite sides of a hardware generation, and which one is real depends on what a 2026 GPU rents for in 2031. The slide that makes the economics work has a five-year contract price in its footnote. Someone did ask about depreciation. The answer was that forecasting is hard.

So that's the next post. I want to rebuild the dollars-per-megawatt model out of the week 3 line items: the $19.2M building, the $40M of IT, the $1.1M a year to run it. Then run payback at a three-year asset life, a five-year one and a seven-year one, and see which of the nine positions survives each case. Asset life is the variable that decides whether the application layer's 7% of AI gross profit is a starting point or a steady state. It decides that for everyone in the stack at once, not only for the people buying GPUs.

All 9 lectures are public on the course site. The guest list is the durable artifact: one practitioner per layer, in order, which is much harder to assemble than it is to record. Thanks to Apoorv Agrawal for running the seminar in the open, and to the nine guests for putting real numbers on real slides. If you model GPU depreciation for a living, whether you're pricing inference, underwriting a data center or just arguing about it at dinner, I'd like to compare notes.