Every lecture in Stanford's CME 296 generates the same picture: a teddy bear reading a book. Noise on the first slide, a finished image on the last. I watched all 8 lectures from Afshine and Shervine Amidi, just over 14 hours of video. The claim underneath them: diffusion, score matching, and flow matching are three coordinate systems on one object. This recap walks the course in order and argues that FID is a regression test the field keeps publishing as a result. One problem won't yield to a better loss, provenance you can strip with a screenshot.
13 Lectures on RLHF: What I Took Away From Nathan Lambert's Course
A reward model reads its score off a single token. One number. For an entire answer. Almost every complaint people have about RLHF (the hedging, the refusals nobody asked for, the model that agrees with whatever you just said) is downstream of that compression. I worked through all 13 lectures of Nathan Lambert’s course and took notes: the canonical recipe, the substitutions that keep replacing its middle stage, and the questions nobody has closed out.
Datacast Episode 133: Full Data Stack Observability with Salma Bakouk
Salma Bakouk is the CEO and co-founder of Sifflet, a Full Data Stack Observability platform. Before Sifflet, Salma was an Executive Director at Goldman Sachs in Sales & Trading in Asia, leading key Data & Analytics initiatives. Salma holds an Engineering Degree from École Centrale Paris in Applied Mathematics and a Master's in Statistics and Data Science.
Datacast Episode 87: Product Experimentation, ML Platforms, and Metrics Store with Nick Handel
Nick Handel is Transform's CEO and Co-Founder. Before Transform, Nick was Head of Data at Branch International in the micro-lending space.
Before Branch, Nick held a variety of roles at Airbnb, both as a Data Scientist & Product Manager. He was on the Growth team that founded the Experiences product and the Data Platform team. His work includes launching Airbnb's ML platform, Zipline, building the company's data science team, helping with the company's initial international expansion, and leading the data science team that launched Airbnb's Trips product.
Before joining Airbnb, Nick was a research economist at BlackRock. He is an avid trail runner, climber, skier, and adventurer, much of the time with his dog Huckleberry.
Datacast Episode 83: Startup Scrappiness, Venture Matchmaking, and Thinking In Bets with Leigh-Marie Braswell
Leigh-Marie Braswell is an investor at Founders Fund.
Before joining Founders Fund, she was an early engineer & the first product manager at Scale AI, where she originally built & later led product development for the LiDAR/3D annotation products, used by many autonomous vehicles, robots, and AR/VR companies as a core step in their machine learning lifecycles. She also has done software development at Blend, machine learning at Google, and quantitative trading at Jane Street.
She is originally from Alabama, graduated from MIT, and loves to play poker, run long distances, and scuba dive.
Datacast Episode 77: Delivering Modern Data Engineering with Einat Orr
Einat Orr is the CEO and Co-founder of Treeverse, the company behind lakeFS, an open-source platform that delivers resilience and manageability to object-storage-based data lakes. She received her PhD. in Mathematics from Tel Aviv University in optimization in graph theory. Einat previously led several engineering organizations, most recently as the CTO at SimilarWeb.
Datacast Episode 66: Monitoring Models in Production with Emeli Dral
Emeli Dral is a Co-founder and CTO at Evidently AI, a startup developing tools to analyze and monitor the performance of machine learning models. Earlier, she co-founded an industrial AI startup and served as the Chief Data Scientist at Yandex Data Factory. She led over 50 applied ML projects for various industries - from banking to manufacturing. Emeli is also a data science lecturer at St. Petersburg State Management School and Harbour.Space University. She is a co-author of the Machine Learning and Data Analysis curriculum at Coursera with over 100,000 students. She also co-founded Data Mining in Action, the largest open data science course in Russia.
Datacast Episode 64: Improving Access to High-Quality Data with Fabiana Clemente
Fabiana Clemente is a Data Scientist with a background that ranges from Business Intelligence to Big Data Development and IoT architecture. Throughout her professional career, she has been leading state-of-the-art projects in global companies and startups. She has an academic background in Applied Maths, and MSc in Data Management combined with nano degrees in Deep Learning and Secure and Private AI.
As YData’s Co-Founder, she combines Data Privacy with Deep Learning as her main field of work and research, with the mission to unlock data with privacy by design. She also aims to inspire more women to follow her steps and join the tech community.
Datacast Episode 47: Math and Machine Learning In Pedestrian Terms with Luis Serrano
Luis Serrano is a Quantum AI Research Scientist at Zapata Computing. He is the author of the book Grokking Machine Learning and maintains a popular YouTube channel to explain machine learning in pedestrian terms. Luis has previously worked in machine learning at Apple and Google, and at Udacity as the head of content for AI and data science. He has a Ph.D. in mathematics from the University of Michigan, a master's and bachelor's from the University of Waterloo, and worked as a postdoctoral researcher in mathematics at the University of Quebec at Montreal.
Meta-Learning Is All You Need
Meta-learning, also known as learning how to learn, has recently emerged as a potential learning paradigm that can learn information from one task and generalize that information to unseen tasks proficiently. During this quarantine time, I started watching lectures on Stanford’s CS 330 class on Deep Multi-Task and Meta Learning taught by the brilliant Chelsea Finn. As a courtesy of her lectures, this blog post attempts to answer these key questions:
Why do we need meta-learning?
How does the math of meta-learning work?
What are the different approaches to design a meta-learning algorithm?









