Poolside 联合创始人详解模型工厂:8周从预训练到发布
Inside the Model Factory — Eiso Kant, Poolside AI
In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines’ recent release nearly 10 times their size.Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside’s Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise, open-source AI, regulation, NVIDIA and TSMC’s influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:How Andrej Karpathy’s RNN work inspired Eiso to start building language models for code in 2015Why Eiso spent four years and $12 million pursuing an idea before the market caredWhy ChatGPT felt like vindication and brought Poolside back to open sourceWhy Eiso would prefer 100 foundation model companies over an oligopoly of fiveThe difference between releasing open weights and publishing genuinely open researchWhy Poolside deliberately built a global research organization outside the Bay Area talent warWhy model building is ultimately 90% engineeringThe Model Factory: Poolside’s end-to-end system for rapidly training and improving modelsHow fewer than 70 researchers run roughly 10,000–20,000 experiments each monthHow Poolside moved from six-month model cycles to five- and eight-week launchesWhy streaming data directly into training unlocked faster experimentationHow immutable data, versioned code, and reproducibility enable rigorous model researchWhy Eiso wants capable researchers to leave their labs and become Poolside’s competitorsWhy 95% of model building can be reduced to better data or compute efficiencyLaguna S and why persistence, verification, and backtracking can outperform raw intelligenceWhy smaller models may handle far more knowledge work than previously expectedWhy reinforcement learning will move earlier into pre-trainingWhy next-token prediction is still failing to extract enough knowledge from the webWhy distillation and environments have become the AI industry’s favorite “drugs”Why mid-training is really an early form of curriculum designLow-precision training, networking bottlenecks, and the next gains in compute efficiencyLaguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launchWhy model builders can often evaluate a new checkpoint within its first 30 minutesModel versus harness: where agent capabilities actually come fromWhy Poolside sees coding and long-horizon software tasks as a path to AGIWhy Eiso thinks MCP and traditional tool calls are “stupid”Why future agents will write scripts instead of choosing from dozens of predefined toolsThe case for minimal harnesses, containers, and model freedomWhy Poolside is prioritizing vision but does not expect to work on audio soonWhy language may be the most compute-efficient modality for encoding knowledge and reasoningThe real cost of model development and why the final training run is anticlimacticThe story behind the Poolside name and why it represents refusing to lower ambitionsHow Poolside raised $500 million while investors still questioned whether AGI was realWhy intelligence could become the world’s most demanded and commoditized resourceWhen open models may become too capable to release without restrictionsWhy unilateral AI safety does not work in a globally competitive environmentHow regulation could accidentally lock in an oligopoly of two or three AI companiesNVIDIA, TSMC, and the hardware systems underpinning foundation-model progressWhy reinforcement-learning wall-clock time is one of Poolside’s biggest bottlenecksWhy Poolside trains models from scratch instead of simply distilling larger modelsHow AI changes the way companies should measure engineering productivityWhy agency may become the most important quality for employees in the AI eraHow leaders align high-agency people through shared goals and clear constraintsHiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside’s Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we’re here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I’m on my way to SF.” I was like, “You’re on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let’s do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don’t have to prepare that much because if you’re truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don’t live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we’re gonna talk about. But I guess, like, what got you into democratization of AI? Like, it’s not obvious from your Li
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力