Context, Compute, Capital, Culture: 13 Lectures From Stanford's CS153

Enrichment is the bottleneck to nuclear. Nuclear is the bottleneck to electricity. Electricity is the bottleneck to AI.

Those three lines sat on a single slide from Scott Nolan (CEO of General Matter). His company is trying to bring uranium enrichment back to the United States, which by his account once held about 86% of the world's capacity. The last American plant closed in 2013. Nolan's was the 7th of 13 sessions. A course about frontier AI spent its midpoint on uranium. That's the whole course on one slide: frontier AI is a stack, and the constraint that binds it keeps moving further down.

Stanford published all 13 sessions of CS153: Frontier Systems, a course taught by Anjney Midha and Michael Abbott (both of AMP PBC). Midha opens with a map of the stack, and each of the other 12 sessions brings in a guest from one layer of it. I worked through all 13 over 3 weeks in September 2026 and took notes. This long-form post walks through them in the order the class met: Midha's opening map, then voice, visual intelligence, unified multimodal models, product management, venture capital, energy, silicon, AI-native companies, data centers, security, scale, and finally the ecosystem Microsoft is trying to build.

What Is A Frontier System?

Midha answered that on the first day with a slide titled CS 153 (and world for next 2 yrs) In a Nutshell. It stacks seven layers from the bottom up: Capital; Land, Power, Shell; Compute Infra; Models & Agents; Applications & Interfaces; Deployment Solutions; and Safety, Policy, Governance. Two columns sit beside the stack, the old system and the new one, under a banner that reads The Great Transition. His point is that every layer is being rebuilt at the same time, because each one has turned into a bottleneck for AI progress. A frontier system is the whole stack, not the model at the top of it. That's why the guest list runs from a uranium company and a venture firm to OpenAI and NVIDIA.

He also gave the class a lens for reading the guests: context, compute, capital and culture. He spent his own hour on the first two. Context is the feedback a deployed model collects about its own work, and whoever has unique access to it captures the value. Compute is scarce and, for now, not interchangeable from one GPU generation to the next. The other two he left to the guests. Ben Horowitz argued that capital has become a real weapon now that GPUs and data can close a lead that hiring engineers never could, and he defined culture as "a set of actions," not a set of beliefs. The guests rarely agreed with each other. They kept arguing about the same four things.

1 - The Map

"Compute, data, algorithms (pretty simple algorithm, transformer) do a little bit of pre training, some fine tuning, good to go. Plug into an app, you got ChatGPT." That's how Anjney Midha (Founder of AMP PBC and a founding investor in Anthropic) described the recipe for manufacturing intelligence 4 years ago, when he and Abbott first taught this class as Security at Scale. His April 2 lecture is about what happened to that recipe: "In four short years we've taken what was a pretty… bespoke process and turned it into an industrial engineering process at scale." Base models now train at least twice a year on roughly 100,000 GB300-equivalents, and the reinforcement learning step alone consumes almost as much compute as the rest of the pipeline combined.

His explanation of RL is the only from-scratch one in the course: "If you had a pet or you've had a sibling who you've had to teach to stay away from your room, you are a successful frontier model trainer applying reinforcement learning." Reward the outcome, withhold the reward when it fails, repeat. For decades that loop plateaued: systems beat people at chess and Go, then stalled. An LLM arrives already knowing a great deal about the world, so RL keeps paying off as compute grows. It's Rich Sutton's The Bitter Lesson with a better starting point.

Of the four bottlenecks, context gets the most time. Whoever has unique access to the feedback a deployed model generates (the repo, the git history, whether the code ran) improves fastest and keeps the customer. His example sits on a slide titled The Context Loop Wars Have Begun: OpenAI agreed to buy Windsurf for about $3B, and within days Anthropic cut Windsurf off from Claude. Midha calls that context leakage, a competitor learning from how your model serves the customers it wants to take. He's an Anthropic investor and says so on a disclosures slide, which is worth remembering here.

The slide I carried through the rest of the course asks, What are the limits to RL? The philosophical answer is that agents should be able to learn anything. The empirical one: "Life is messy. Progress is fastest in easily verifiable domains." Midha takes the second. Code has tests. Materials science has experiments. Aesthetics don't.

Compute is where the lecture turns into prices. By AMP's own tracker, an H100 rented for $1.73 an hour two years ago, and it costs more now on an older chip. The morning of the lecture, a founder with somewhere between $700M and $1B raised messaged him: "Take them right now. Price not a problem." Hyperscaler capex ran about $300B last year, $600B this year and a planned $1.2T next year, by his numbers. His case that this isn't madness is a ratio. Land, power and buildings trade at 3–4x revenue. Software revenue trades at 30–40x. So a dollar of steel and power that becomes a dollar of software revenue is worth about 10x what went in. The catch is that compute isn't a commodity yet. It has no common unit, no standard interface, no pooling and no way to swap one supplier's GPUs for another's. Electricity and networks got there through standards like AC/DC and TCP/IP, and institutions to enforce them.

He ended on homework. "What will it take to ensure a peaceful transition on compute?" In June, OpenAI published Industrial Policy for the Intelligence Age, which argues that "incremental policy updates won't be enough." That's one lab's answer. Midha asked 500 students for theirs.

2 - The Future Of Voice Systems

In late 2021, Piotr Dąbkowski was about to watch a movie with his girlfriend, who didn't speak English, so they switched it to Polish. What came out was the voice he and Mati Staniszewski (Co-Founder and CEO of ElevenLabs) grew up with in Poland: one flat narrator reading every character, men and women alike. The two had been friends since high school in Warsaw. The company they started was supposed to fix dubbing, and dubbing takes three models: transcription, translation and speech. Customer interviews narrowed the first product to the third. Creators didn't need a dubbed film yet. They needed to fix a bad line or record a voice-over without being on camera.

Speech was also where the room was. "There was very little research done in audio. Most people focused on LLMs," he told Sequoia's Training Data. Audio data is scarce and rarely comes with an accurate transcript, let alone a label for how a line was said. His example is "What a wonderful day," which flips meaning when it's said sarcastically. The team built on James Betker's open-source Tortoise model, trained the first checkpoint on what he calls "a tiny amount" of compute, in the tens of thousands of dollars, and skipped a roughly $6,000 patent because the field was moving too fast for one to matter. Voice cloning and a voice marketplace followed in 2023, working dubbing in 2024, and real-time agents in 2025.

The technical question he spends the most time on is how to build a voice agent. A cascaded agent runs three models in a row: speech to text, a language model, then text to speech. A fused model goes from audio to audio in one pass. Fused is faster, around 300 milliseconds by his account, and more expressive. Cascaded is easier to inspect, to guard and to connect to tools. His example is rebooking a flight, which means checking identity, taking payment and calling several systems in the right order. For that, he expects enterprises to stay with cascades "for the next few years," and perhaps to use a hybrid that talks in fused mode and hands off to a cascade when money moves. It's Midha's point about verification in engineering form: every arrow in a cascade is a place to check the work. Reliability beats latency when the agent can move money.

The business is where the numbers get large. ElevenLabs closed 2025 at $330M in revenue and added more than $100M of ARR in its biggest quarter, putting it above $430M with more than 450 people, most of them in teams of fewer than 10. Roughly half the revenue is enterprise and half self-serve. His pricing rule is to "work backwards" from the value delivered and capture about "one tenth" of it. He's also blunt about the risks his own product creates. Voice authentication for banking, he says, "is not the future." One charity project routed likely scammers to an ElevenLabs agent built only to waste their time. And the company has restored the voices of almost 10,000 people who lost them to ALS or throat cancer.

A week later, Amit Jain would argue for the opposite architecture: one model doing everything at once.

3 - Visual Intelligence

Training a strong diffusion model in pixel space "often consumes hundreds of GPU days." That line comes from the abstract of High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2022), and Andreas Blattmann (Co-Founder of Black Forest Labs) is its second author. His university lab was up against Google and OpenAI, so it couldn't spend hundreds of GPU days. The fix was to compress first: train an autoencoder that keeps what people notice in an image, then run diffusion in that smaller latent space. Midha calls it a learned JPEG. That work became Stable Diffusion in 2022.

Blattmann's thesis in the interview is that intelligence should be learned from natural signals like video and audio, the way people watch before they act, and not from language with perception bolted on. His own papers trace that path. Align Your Latents (CVPR 2023) turned an image model into a video model by adding a time dimension to the latent space, and in doing so turned Stable Diffusion into text-to-video at up to 1280×2048. Stable Video Diffusion (2023) split training into three stages: text-to-image pretraining, video pretraining and high-quality fine-tuning. Its main finding was about data: good video needs a well-curated pretraining set, and the paper spells out how to build one.

Black Forest Labs didn't start with all of that. Image models still drew hands with the wrong number of fingers, so the team aimed for a model "ten x better than everything else." FLUX.1 found fit before its API was public. Then users told them what came next by training LoRAs to keep characters consistent. That became FLUX.1 Kontext in May 2025, which takes text and images as the prompt and edits in place, up to 8x faster than GPT-Image by BFL's own benchmark. Midha tells the story behind it. A rival editing model appeared, the team reorganized within 24 hours, Kontext shipped about 60 days later, and its revenue doubled within six weeks. In June, Kontext [dev] released the 12B editing model as open weights that run on consumer hardware. The lineup splits by license: Schnell is Apache 2.0, Dev is open but paid for commercial use, and Pro is served by API.

The business followed. In September 2025, Bloomberg reported that Meta would pay $140M over several years to use BFL's image models. Midha says BFL was about 25 people when it took on image editing for Meta's roughly two billion users, and that it now has "hundreds of millions in revenue." He also sits on BFL's board, so those are an insider's numbers.

The part I found most useful is Blattmann on verification. A robot arm either completes a motion or it doesn't. Taste in images varies by audience. Midha's argument for open weights follows from that: where preferences differ, let users tune the last mile themselves. BFL's Self-Flow work, published in March 2026, attacks a different limit. Diffusion models usually borrow their representations from a frozen encoder like DINOv2, and the generator can only get as good as that teacher. Self-Flow has the model learn its own representations across images, video and audio. VentureBeat reports it converging 2.8x faster than REPA, the standard alignment method.

One question stays open. Blattmann doubts explicit 3D is the right interface for spatial intelligence, because people learn space from video, sound and moving through the world. Midha thinks point clouds and meshes still earn their place in robotics. They didn't settle it.

4 - Unified Intelligence Systems

The slides for this session were the demo. Amit Jain (Co-Founder and CEO of Luma AI) sketched a mind map, handed Luma's model one example slide, and let it generate the deck. He deleted one version he didn't like and kept the rest. The point of the exercise is his thesis: the next useful AI is one model that reads and generates text, images, video, audio and code in a shared space, not a language model with image and video models attached. He put the case most directly on Amplify Partners' podcast: "Unless you're able to jointly train and backprop through, it's very difficult to imagine how the sum of the parts is better than the parts themselves."

He got there by following the data. Jain worked on LiDAR, the Titan car project and Vision Pro at Apple, and Luma started as a 3D capture app built on NeRFs and Gaussian splats. In 2023 he was describing a future where "someone with an iPhone can now do this maybe 80% of the way." But an app can't collect data at internet scale, and the internet is full of video. So Luma moved there. Dream Machine reached about 6M users in its first 3–4 weeks, and its successor was trained on "1.2 quintillion tokens," by Jain's count. The user signal turned out to be noisy: some people downloaded bad outputs to show how bad AI video was. The lesson he draws is to design the algorithm around where the data already lives.

The factory behind Luma has about 30 PB of trainable multimodal data and roughly 10,000 GPUs. What it ships has three layers. Skills are human-written knowledge, like the 50-page slide-design guide behind this talk. The tool harness handles Linux, APIs, OCR and deployment. The unified model sits on top and decides what to do. That product launched in March 2026 as Luma Agents, running on Uni-1, a decoder-only transformer that interleaves language and image tokens. Luma has raised about $1.5B, and Jain says Coca-Cola is moving $3B a year of content production to it. Midha sits on Luma's board.

Then the prediction: these systems overtake language-only models within 1–3 years, because they can learn from far more kinds of data.

This is the claim in the course I disagree with most, and the course supplies the reasons. The case for it is real. More data types is real leverage, the slide demo worked, and a brand moving $3B of production is real money. But three earlier sessions point the other way. Midha's first lecture says RL progress is fastest where work can be verified. Blattmann says verification is the bottleneck for images, because taste varies by audience. Staniszewski keeps enterprises on cascades because every step can be inspected and every tool call checked. Put those together and the unified model's advantage, doing everything in one pass, is exactly what makes its work hard to check. So the course's own framework predicts unified models win first where nobody has to verify the output, in creative work judged by taste, which is where Luma's named customers are. It predicts they win last where the work has to be checked.

Luma's own product agrees. By LBB's account, the agents route each step to the best model for it, run quality checks, reject weak assets and try again. That's a layer of checks around the unified model, not one pass. Jain conceded the rest on The Cognitive Revolution: individual capabilities like object recognition "are still worse than dedicated computer vision models." My position is that the course's boldest prediction is contradicted by its first lecture. Unified models will overtake language models on breadth long before they overtake them on the work people pay to have checked.

One example doesn't fit, and I can't wave it away. Jain described an energy customer whose unified model read grid diagrams and grid code and produced schematics that coding models couldn't, because those models couldn't read the layout. That output can be checked against the code. It's a verifiable domain, and the unified model is already winning it.

5 - Product Management In The AI Era

In July, Nikhyl Singhal (Founder of Skip, formerly a product executive at Google and Meta) wrote about offers to product executives reaching $10M a year. "The packages look less like salary bands and more like pro-athlete contracts," he wrote in The Skip, his newsletter with more than 25,000 subscribers. Three months earlier, in this session with Michael Abbott, he gave both halves of the market at once. Top PM pay has more than doubled in 18 months. Some large tech companies are discussing layoffs of 30–70%. Both are true, and the session is his explanation of why.

It starts with a definition. PMs are the connective tissue between the people who build and the people who sell. Abbott frames the history: enterprise PMs wrote a PRD and handed it to engineering, Apple built through designers and engineers without a classic PM layer, and now designers can vibe code. Singhal's estimate is that about 80% of PM work is moving and packaging information: status decks, the "dog and pony show" reviews assembled through layers of management. AI can now read the support chats, sales calls and surveys directly and rank them by revenue, complexity and fit with the brand. What's left is judgment about what to build. So the information movers get automated and the product builders get paid. The most exposed, in his view, are mid-career managers who were promoted to coordinate. The 80% is his estimate, not a measurement.

His map for the role is an S-curve. Before product-market fit, founders take as many shots on goal as they can and there's no PM function to speak of. Once customers start pulling, the company needs consistency and process. Hypergrowth brings the big PM organizations. Late-stage companies have to beat the innovator's dilemma to build anything new. By his numbers, perhaps 1–4% of companies find fit and 1–2% of those reach hypergrowth. The lesson he draws from Google is speed. Hangouts tried to solve Google's internal problem of unifying Gmail, Android and voice, while WhatsApp made plain messaging reliable first. Chrome shipped every 6 weeks, Firefox every quarter, Internet Explorer once a year.

The advice is about careers, which is Skip's business. A 40–50 year career at 2–3 years per company comes to 15–18 jobs, so he treats a career as "chapters in a book," not periods in a hockey game. His list is practical: stay hands-on with current tools, build judgment about how systems fit together, keep a network, join a company growing faster than you are, and treat comfort as a sign it's time to move. Skip itself is an invitation-only community of about 125 heads of product, including leaders from Anthropic, OpenAI and Meta.

One thing in his argument doesn't close. The prescription for the exposed middle is reinvention, and The Skip ran a whole post in April calling it "the product skill you must now master." But his own career framework from November 2025 says the opposite about who can make the jump: "A class isn't going to make you into a Builder. You need to have the obsession." If being a builder takes an obsession no course can supply, the coordinators he says are most at risk can't reinvent their way out. He argues both, and I don't know which one he believes more.

6 - Venture Capital Systems, Network Effects

In 2009, the working assumption in venture capital was that about 15 technology companies a year would ever reach $100M in revenue. Ben Horowitz (Co-Founder of Andreessen Horowitz) and Marc Andreessen thought software would push that number toward 200, and they built a firm for the bigger number. Anjney Midha, who worked for Horowitz at a16z for several years, introduced him on April 21 as "the Quincy Jones of technology." The session is about treating a venture firm as a system you can redesign, and about what changes in that system when AI arrives.

Most of the design choices came from one observation: traditional VC was a good product for LPs and a poor one for founders. So a16z shared economics but centralized control, because a partnership where everyone shares control can't reorganize; whoever would lose power blocks the change. That structure is what let the firm move into American dynamism, crypto and bio. It split into small groups, because nobody has a real investing conversation with 30 people, and it spent management fees on relationships instead of salaries. One trick drew on Horowitz's earlier sale to HP: a16z called HP's enterprise briefing center every week to learn which companies were visiting, then invited them over to meet its startups. He owns up to starting the feud with other VCs, including a blog post titled "Four Things that VCs Do That I Don't Like" and a Lil Wayne line about seeing "the trigger and the middle finger" when rivals came at him with a peace sign. His read is that the hostility kept competitors from copying the model.

The AI half of the talk is Midha's capital bottleneck, argued from the other side of the table. In the old software world, hiring 1,000 engineers couldn't close a two-year lead, because the work wouldn't split that many ways. With enough GPUs and data, many problems now yield to money. Code and UI are weak moats. Demand is running hot enough that he jokes, "you mean companies didn't always go from nine to thirty billion in run rate in like six weeks." Culture, the fourth bottleneck, gets his cleanest line, borrowed from Bushido: a culture "is not a set of beliefs; it's a set of actions."

The Q&A's best story is Slack. Stewart Butterfield turned a failing game into Slack with about $6M left, which is Horowitz's evidence that "the one unforgivable sin" is running out of money. He also rejects the "SaaS apocalypse," pointing to Navan's airline and hotel supply relationships as a moat no model lab will bother to rebuild.

His message to the students is that disruption favors them. They'll learn the new tools before the people above them, and the claim that companies will run on bots with no employees is overstated. The early data isn't that kind to them yet. Stanford's Digital Economy Lab launched its AI Economic Indicators in June. Its Canaries dashboard, built with ADP, shows employment for 22-to-25-year-olds in the most AI-exposed occupations declining since ChatGPT, while the less exposed groups grow. Its July reading on the wider economy: "We see no decisive evidence of transformation at present." The Forecasting Research Institute survey says much the same about the future. Its economists' unconditional forecasts are "close to historical trends," and they give rapid AI progress by 2030 about a 14% chance. In that rapid case, labor force participation falls from 62% to 55% by 2050.

So the revenue curves in the room and the payroll data outside it don't match. Horowitz may be right that the tools favor the young. The first measurements say entry-level jobs in the most exposed fields are where the losses show up first.

7 - Energy Bottlenecks

The Paducah Gaseous Diffusion Plant in western Kentucky started enriching uranium in 1952, first for weapons and later for power plants. It stopped in 2013. In August 2025, General Matter broke ground on the same property, and on January 7, 2026, the Department of Energy awarded it a $900M, 10-year, milestone-based contract to build enrichment capacity there. It was one of three $900M awards in a $2.7B package, alongside American Centrifuge Operating and Orano. Scott Nolan (CEO of General Matter, formerly of SpaceX and Founders Fund) had started on the problem in December 2022 and founded the company in January 2024. The contract came about 24 months later.

His session with Midha on April 23 is the bottleneck chain from the top of this post, taken one link at a time. The first wave of AI data centers ran on stranded power: hydro, wind and flared gas, a pattern Bitcoin miners found first and Crusoe carried into AI. That supply is used up, and only new generation can meet what's coming. Data centers want high uptime, so developers are buying gas turbines, which are scarce, while they wait for baseload power. Nuclear is his answer. His chart puts its 2024 capacity factor at 92.3%, the highest of the sources shown, and Google, Microsoft and Meta are all looking at it.

Nuclear fuel moves through five steps: mining and milling, conversion, enrichment, deconversion and fabrication. Reactors need fresh fuel every 1 to 10 years, depending on the design. Enrichment is the step the U.S. gave up. It held about 86% of world capacity during the Cold War. After the Megatons to Megawatts program and years of imports from Russia and Europe, the plants stopped making money and closed. General Matter's target is HALEU, high-assay low-enriched uranium, the fuel many next-generation reactor designs need and today's plants don't. On Nolan's slide, enrichment is the largest single cost in HALEU fuel. The company expects to put about $1.5B of private money into the Paducah site, hire 140 people and be running before 2030. "This award is really going to help us build at a larger scale more quickly," he told WKMS.

He's candid about the calendar. The next few years are the hard ones, because large reactors take 5 to 10 years to build and industrial power equipment is already scarce. He expects "tens, not hundreds" of small modular reactors soon and the gigawatt-scale buildout in the early 2030s. When the conversation turns to space, he'd bet on SpaceX for orbital data centers, since no one else launches at its volume. His career advice is the same shape as his company: ignore the hype and the panic, find an important unsolved problem, and work where you're most useful.

The chain has the order right. What it doesn't have is a single clock. Turbines now. Reactors in the 2030s. Fuel before the reactors, if General Matter hits its date. On the slide, enrichment is the bottleneck to AI, and Nolan says it holds "on a five year time frame." By his own schedule, though, the reactors that would burn HALEU arrive at about the same time as the fuel. Either one could end up being the link that binds, and nobody in the room could say which.

8 - The Compute Behind Intelligence

"The ChatGPT moment for physical AI is here." Jensen Huang (Founder and CEO of NVIDIA) said that in January, when NVIDIA released Alpamayo 1, a 10B-parameter open model for self-driving cars that turns video into a trajectory plus a written trace of why it chose it. When he sat down with Midha on April 30, a self-driving car was his example of why computing is being reinvented for the first time in more than 60 years. A car can't look up its next move. It has to generate one, continuously, from a scene nobody recorded before. Most computing used to retrieve stored answers. Now it produces new ones all day, and agents run for hours with tools and context instead of answering one prompt.

Transistor scaling can't carry that alone, so NVIDIA's answer is co-design: algorithms, compilers, chips, interconnect, software and data centers designed together. Midha proposed tokens per watt as the metric that matters, and Huang agreed, especially for inference. Peak FLOPS is the wrong number to stare at. When a model decodes, it spends most of its time moving weights and state through memory, so aggregate memory bandwidth decides the speed. Prefill and decode behave differently enough that they're worth separating.

His sharpest technical point is that model FLOP utilization, or MFU, is a misleading target. The reason is Amdahl's Law: if a fraction of a job can't be sped up, that fraction caps the whole job. If 5% of a training step is stuck waiting on one slow component, no number of extra GPUs gets you past a 20x speedup. So the rational design is to over-provision the slow parts, even if some expensive chips sit idle.

The roadmap follows the workload. Hopper was built for pre-training, NVLink 72 for bandwidth-heavy inference, and Vera Rubin for agents, which need storage on the fabric and CPUs with very low latency. The public numbers show the pace. TechCrunch reported Vera Rubin at up to 50 petaflops of inference per package in the second half of 2026, against 20 for Blackwell Ultra, with Rubin Ultra at 100 in 2027 and Feynman in 2028. One roadmap summary puts a Rubin Ultra rack at about 600 kW. A new architecture every year, and rack power keeps climbing.

He also argues AI is bigger than language, and NVIDIA ships models to prove it: Nemotron for language, BioNeMo for biology, Alpamayo for driving and Groot for robotics. He backs open models because a black box is hard to inspect or secure. Asked whether compute is scarce, he calls the question incomplete. Scarce for what task, at what performance? Cheaper intelligence keeps creating new uses, so there's no fixed amount of compute that would be enough.

That's the logic behind the biggest number in the session: computing could need about 1,000x today's energy, and he says that estimate could be off by a couple of orders of magnitude. NVIDIA's own scaling-laws post supplies part of the mechanism: test-time reasoning "can easily require over 100x compute for challenging queries." Every generation raises tokens per watt, and he still expects total energy to climb 1,000x. The claim is that demand will outrun efficiency by orders of magnitude. Scott Nolan's reactors are the supply side of that bet, and they arrive in the 2030s.

The Q&A gets personal. Huang rejected the comparison of GPUs to atomic bombs, since a general-purpose technology is defined by what people build on it. He named NVIDIA's biggest business mistake as shifting resources to mobile, then being shut out by Qualcomm in the move from 3G to 4G. His method under uncertainty is to reason from what he can observe, build a model of the future and work backward from it while keeping options open. Midha also got his Denny's order out of him.

9 - The AI-Native Company

The README for gstack, Garry Tan's open-source Claude Code setup, opens with a productivity claim: 11,417 logical lines of code a day in 2026, against 14 a day in 2013. He also counts 3 production services and 40+ features shipped in 60 days while running Y Combinator full-time. At the time of writing, the repo has more than 130k GitHub stars. Tan (President and CEO of Y Combinator) and Diana Hu (General Partner at YC) gave the May 5 session, titled "The AI native company: How one founder becomes a 1000x engineer." Their claim is that the unit of production has changed to "human + agents + memory + evals + customer loop." In 2010, a 10-person startup felt lean. In 2026, their slide says, "a 6-person team can hit $10 million in revenue."

Tan's half is about how one person builds. He grants the usual complaints about AI coding: slop, hallucinations, demos that aren't production, lines of code as a bad metric. His answer is structure, not less AI. Move from a copilot to a software factory. gstack is his factory, 23 slash-command tools that play CEO, designer, eng manager, release manager, doc writer and QA. /office-hours is modeled on YC partner sessions and asks 6 forcing questions before any code gets written. /plan-ceo-review challenges scope, and /retro closes the week. The skills are runbooks written in Markdown, and resolvers in an AGENTS.md-style file decide which one runs, so the agent doesn't load everything into context at once. Where exactness matters, he writes code instead of trusting the model. His example is TypeScript that hands the agent the current time and upcoming events as JSON. GBrain, his memory system, stores hunches and links in a repo the agents can search.

Hu's half is about the company around that person. In an open-loop company, information lives in heads, DMs and unrecorded meetings, and agents see about 10% of the company's state. A closed-loop company makes every workflow leave something agents can read: Linear tickets, GitHub commits, Slack channels instead of DMs, recorded sales calls. Her slide says that gets you half the sprint time and 10x the shipped output. Her line on evals is the best in the session: "Evals are taste made executable." Founders have to write them, because founders hold the quality bar. On strategy, "the best AI startups don't demo intelligence. They deploy solutions." Salient closed top U.S. banks for loan servicing through a forward-deployed pilot. HappyRobot's logistics voice agents grew revenue 10x in under a year. She suggests taking a job inside the target industry before writing any code.

Then the number I'd keep from this session. Across the last three YC batches, 10% week-over-week revenue growth is the average. Five years ago it was the top 1%. Their closing slide borrows this course's own tagline: "CS 153 calls itself the one-person frontier lab. Our message: That lab can become a company."

One thing about the README bothers me. Tan's own slide lists lines of code as a bad productivity metric, and the first number his repo offers is lines of code per day, even after he strips out inflated raw counts. He says the point "isn't who typed it, it's what shipped," and I agree with him. His headline number just doesn't measure that.

The bigger question this session raises is one the course never answered. Two weeks earlier, Ben Horowitz said capital has become a real weapon, because GPUs and data can now close a lead that hiring engineers never could. Here, YC says a 6-person team can reach $10M and the median company in a batch grows 10% a week. Back in the third session, Midha described Black Forest Labs taking on image editing for Meta's two billion users with about 25 people. Same course, a few weeks apart. Whether AI favors the biggest balance sheet or the smallest team, nobody on stage said.

10 - The Discipline Of Delivering Value Per Gigawatt

A TPU rack at Google holds 64 chips. Between racks, the fiber runs into optical circuit switches, chips covered in tiny mirrors that tilt to steer light from one fiber into another. When a TPU fails, software pulls its whole rack out of the network and swings a healthy one into the exact same position. In Amin Vahdat's words, "the torus becomes whole again," and it takes seconds. Vahdat (SVP and Chief Technologist, AI & Infrastructure at Google) sat down with Midha on May 7, and that swap is his whole argument in one mechanism. Nothing new gets built. The same hardware does more useful work.

He thinks the build-out is measured wrong. Google is aiming for tens of gigawatts over the next four years, and Midha's back-of-the-envelope math puts a gigawatt at about $40B of infrastructure. The number Vahdat cares about is value per dollar, per watt and per gigawatt: revenue, users, capability, not how many gigawatts you own. Twice the value per gigawatt means half the gigawatts. This is Midha's "compute" C from the operator's side, and it meets Huang's tokens per watt from the other end. Huang measures from the chip. Vahdat measures from the building.

Reliability is his first example of where value leaks. Five nines of availability takes one-plus-one redundant power feeds, so half the power capacity sits unused at any moment. So Google now asks customers to choose: twice the capacity with a few days of downtime a year, or half the capacity with almost none. Enterprises still pick five nines. The frontier labs say, "Give me more capacity. I'll take the downtime." Training a frontier model is about throughput. Amdahl's Law shows up again, same as in section 8: accelerators sit idle when HBM, host CPUs, storage or the network can't feed them, and mixture-of-experts models make that worse.

Topology follows the workload. The mirrors build 3D tori, which fit all-reduce, the collective that dominated training. A switched fabric handles all-to-all traffic better, so Vahdat calls the optics "an augment" to Google's electrical packet switches, not a replacement. Lumentum, which sells these switches, explains the appeal: a MEMS switch "consumes energy primarily during reconfiguration of mirrors." Vahdat also told on himself. In the TPU v2 design debates, he argued for Ethernet as the interconnect, and he was wrong.

The binding constraints are physical: memory supply, land, permits and a grid that's largely spoken for. Utilities will promise five or ten gigawatts, he says, if you agree to pay for all of it for the next 20 years. Google depreciates compute hardware over six years, and older chips stay busy because demand is that high. In the long run, energy is the bottleneck. He says the US is underinvested in solar, wind, batteries and the factories that make them, and he treats space data centers as worth exploring alongside those, not instead of them.

Here's where I'd push. The downtime trade is a training deal. Vahdat says so himself: an enterprise-grade service is five nines. And the build is tilting toward serving. AFL's white paper Architecting AI at Scale puts inference at 50–70% of total AI compute demand. A May preprint, Separating Intelligence from Inference, proposes the escape route: train centrally, run inference on devices near the user, and save an estimated 19 TWh a year at one billion daily users. I don't buy its premise that inference "does not architecturally require centralization." Agents carry long context and shared state, and AFL's agentic pods need memory access in under 15 microseconds. That's a data-center number, not a phone number. My bet is that inference stays in the building and inherits the five nines. Value per gigawatt is still the right metric. It just gets harder to double once most of the gigawatt serves customers who won't take the downtime.

11 - The Road Ahead: Resilience Required

"If you find a vulnerability, please tell us about it. We promise we won't sue you." That was PayPal's responsible-disclosure policy in 2007, and Joe Sullivan (Founder of Joe Sullivan Security, CEO of Ukraine Friends and Venture Partner at Costanoa Ventures) wrote it. He went from DOJ prosecutor to running security at eBay and PayPal, Facebook, Uber and Cloudflare. At Facebook, his first reaction to hackers asking for money was a prosecutor's: "How can I use the law against you?" His team told him to shut up and pay them. In 2010 or 2011, Facebook launched what he calls the third bug bounty anywhere. His May 14 talk is the only lecture in the course built on a personal failure.

That failure is the 2016 Uber breach, which exposed data on 57 million riders and drivers. Uber treated the people who emailed about the stolen database as bug-bounty participants, paid them $100,000 and had them sign NDAs. The advice he'd always gotten worked like trespass law: invite someone in after they step into your yard, and it's no longer trespass. At trial, the jury was instructed that Uber couldn't give that permission. "It just basically gutted our whole defense," he said. He was convicted of obstruction and misprision of a felony, then sentenced in May 2023 to three years' probation and a $50,000 fine. The Ninth Circuit upheld the conviction in March 2025, holding that "an actor's authorization, or lack thereof, is assessed at the moment of access."

He expected the security community to shun him. Instead, more than 200 people wrote letters of support, and his peers at a Black Hat CSO summit gave him a standing ovation. Michael Abbott asked what he'd do differently at Uber. "I think everything we did, I would do the same. I wish we had more documentation." He'd also build trust with the other executives before a crisis, not during one. Document your decisions before you need them.

Then he widened the lens. Ransomware turns a breach into a shutdown. Jaguar Land Rover was hit in late August 2025, and Sullivan remembered production stopping for three months. Public reporting puts the full stop at about five weeks, backed by a £1.5B UK government loan guarantee. The UK's Cyber Monitoring Centre estimated the cost at £1.9B ($2.5B) across more than 5,000 organizations, "by some distance, the single most financially damaging cyber event to ever hit the U.K." A week earlier, Vahdat's frontier labs had volunteered for a few days of downtime a year. A carmaker with thousands of suppliers can't make that trade. Resilience now means keeping the business running, not only protecting the data.

AI makes that harder. One bank he works with went from about 250,000 lines of code a month to about 1.25 million. Much of the new code comes from people who aren't engineers. Garry Tan's 11,417 lines a day, from section 9, is Sullivan's problem seen from the other side. No single scanner covers that volume, so he argues for layered monitoring and guardrails around the tools.

12 - Scale, AGI, And The Future Of Everything

"Since we can't figure out what product to build, we're just going to put this into an API." That's how Sam Altman (CEO of OpenAI) described the GPT-3 launch when he sat down with Midha on May 21, back at Stanford where he once taught How to Start a Startup. The API got "no traction at all" at first. His theme for the session was that people misjudge exponential progress, and they often miss a new capability until a product makes it concrete. OpenAI was on that list too.

The three GPT papers show how early the capability was there. GPT-1 (2018) pre-trained a 12-layer Transformer on more than 7,000 unpublished books, then fine-tuned it per task, and beat the state of the art on 9 of 12 benchmarks. GPT-2 (2019) scaled to 1.5B parameters on WebText, 8 million pages linked from Reddit, and set the state of the art on 7 of 8 language-modeling datasets zero-shot. The paper also noted that the model "still underfits WebText," which was an invitation to go bigger. GPT-3 (May 2020) did: 175B parameters, "10x more than any previous non-sparse language model," with tasks specified in the prompt and "no gradient updates or fine-tuning." The capability was on arXiv in May 2020. ChatGPT shipped in November 2022. Two and a half years.

What closed the gap was a chat box. Developers who couldn't make the API work for their businesses were using their keys to talk to the model. "We can build a good chat box. People clearly want that," Altman said. Word of mouth did the rest, and OpenAI scrambled to build a company around what it realized was a killer product. Coding followed the same path: the plan before ChatGPT was to go all in on code, because code is how models will control computers. He put the turn for Codex at GPT-5.5. The capability shows up before anyone notices. The product makes people notice.

His analogy for where this lands is electricity. Early power companies didn't sell electricity, because nobody knew what it was or why they'd want it. They sold "light at night." "We are in the process of creating a new utility," he said, and people will buy tokens the way they buy light, without thinking about the hardware. Asked what he'd build as a one-person frontier lab, he picked inference: "I think we have not invested enough in being able to deliver at scale huge amounts of cheap intelligence." He expects the frontier labs to turn into inference companies. That's the same shift Vahdat described in section 10, seen from the model side.

In the Q&A, he rejected Yann LeCun's claim that LLMs are a dead end. The models have "far surpassed human intelligence in some ways," he said, and are "wildly worse" at long-horizon, high-judgment work. He blamed the field's slow start on "a generation of scientists who just were way too certain" about what wouldn't work. He also worried that schools teaching as if AGI weren't coming would lead to an "atrophy of learning how to think." On compute, he argued the shortage may never end, since demand for energy looks completely different if the price falls 10x. Every price drop pulls in more demand than new supply covers. Midha's H100 prices from section 1, higher now than two years ago on an older chip, are what that looks like in practice.

I'd push back on one part of his framing. The skeptics weren't the only ones who misread the exponential. OpenAI read it too, from the inside, and still couldn't see a product in a 175B model that wrote convincing news articles. The GPT-3 paper measured translation, question answering and cloze tasks. Nobody's benchmark measured whether people would want to talk to it. The misjudgment was never about believing in scale. It was about not knowing what scale was for until someone used it.

13 - Building The Frontier Ecosystem

Microsoft's newest Copilot, Scout, is "a little bit of an autopilot." It runs with a heartbeat, takes your Entra ID and acts on your behalf as a long-running agent. It writes and runs its own code, so Satya Nadella (Chairman and CEO of Microsoft) spent part of the final session explaining how to keep it contained. Microsoft is working with the OpenClaw foundation and running agents inside a new sandbox container, which the captions render as MXC. Scout is the third Copilot form factor, after chat and Cowork. Michael Abbott hosted, and he opened where most people would: the 2019 bet on OpenAI.

Nadella called that bet the product of a "prepared mind," years of searching for a breakthrough in natural language before scaling laws produced "stunning results." Inside Microsoft, "there was no uprising." The real question was "how do you allocate scarce resources, particularly compute, to specific efforts." Seven years later, the scarce resource is the same, and his answer is an ecosystem.

His central claim is that every company should run its own hill-climbing machine: a good model placed inside an environment the company builds, where it "can go learn using the traces of that company and those tasks." The company writes the RL environments and the private evals, welcomes frontier, open-weight or licensed models into that gym, and keeps the IP it produces. He calls that the only way the AI economy stays "positive-sum," with many participants at the frontier instead of a few model owners renting to everyone else. Firms that can't build the machinery get Microsoft's tooling: "The easy button is there." This is Midha's context loop from section 1 in a CEO's vocabulary. The course ended where it started.

The hardware half argues for designing the whole system rather than buying more GPUs. Microsoft announced Maia 200 in January: a TSMC 3nm inference chip with 216GB of HBM3e, which Microsoft says delivers 30% better performance per dollar than the hardware already in its fleet. Nadella said it's co-designed with Microsoft's own models and OpenAI's, and it runs alongside Cobalt ARM CPUs tuned for agentic workloads. That's Vahdat's value-per-gigawatt argument from the Azure side. He also covered devices, including a Dev Box with a petaflop of AI compute and 128 GB of unified memory. On quantum, he was blunt: "Quantum is not going to replace classical." It's a new accelerator for simulating nature, married to classical machines that handle storage and memory.

Here's where I'd push. The hill-climbing machine is the right answer to section 1's context war. It lets the company own the loop that improves the model on its work. The easy button weakens that. If Microsoft runs the machine, the traces and evals pass through a company that also trains its own MAI models and licenses OpenAI's. Nadella's answer is to guard those assets the way firms already guard confidential data. Contracts are good at protecting data. Midha's Windsurf story wasn't about stolen data, though. It was about a model owner learning how its model serves the customers it wants. Contracts police that poorly. I'd rent the frontier model and own the loop, and I'd be slow to hand the loop to anyone who also sells models.

One recap of Microsoft's Build 2026 keynote sums up Nadella's pitch the same way: developers should build a "frontier intelligence ecosystem" that they own and control. Owning it is the hard part.

Last Takeaway

The number from this course I'll carry into the next post is six. Google depreciates its compute hardware over six years, Vahdat said, and older chips stay busy because demand is that high. In my MS&E 435 recap, I complained that the number deciding whether the AI build-out pays, how long a GPU keeps earning, never made it onto a slide. Here an operator said it out loud. It doesn't settle the question. Six years is one company's accounting policy, not a rental price in year six, and it sits next to Huang's roadmap in section 8, with a new NVIDIA architecture every year. A six-year schedule on hardware that gets a successor every year is either a bet that demand never lets the old chips go idle, or a writedown nobody has taken yet.

This post joins my recaps of CS 329A on self-improving agents, CME 296 on diffusion models and MS&E 435 on the economics of the AI stack. Each covered its own layer. Read together, they keep circling the same two questions.

  • Who checks the work: CS 329A's argument was that the generator has raced ahead of the verifier. Here, Midha said progress is fastest in easily verifiable domains, Diana Hu called evals "taste made executable," and Nadella built his hill-climbing machine on private evals. CME 296 spent a whole lecture on evaluation, and I came out of it arguing that FID is a regression test the field keeps publishing as a result.

  • How long the hardware earns: MS&E 435 priced a megawatt at $59M and never priced its asset life. CS153 gave me one real figure and one reason to doubt it.

That's the next post in this series. I want to take those two questions across all four courses, line up what each guest actually claimed, and see which claims survive contact with the other three syllabi. Stay tuned for it.

If you want the source material, start with the course site and the YouTube playlist. One warning before you press play: the lecture numbers run backwards against the calendar, so Lecture 1 is Nadella's session and Lecture 13 is Midha's opener. This post follows the order the class met. Thanks to Anjney Midha and Michael Abbott for putting a uranium company, a venture firm, OpenAI and NVIDIA on the same syllabus, and for publishing all 13 sessions where someone outside the room could take notes.

If you're pricing GPUs on a depreciation schedule, or writing the evals your company hill-climbs on, I'd like to compare notes before I write the synthesis.