Python Polars v2.0.0:破坏性变更与弃用 cut/qcut
Python Polars 2.0.0
推荐理由
v2.0.0 包含多项破坏性变更(如 SQL 窗口函数行为改变、ENUM 类型解析变化)及核心 API 弃用,直接打断现有用法。请检查代码中是否依赖旧版 SQL 语义或使用了已弃用的 cut/qcut,并尽快适配。
💥 Breaking changes
💥 破坏性变更
- Run SQL window functions over grouped rows on the aggregated rows, and evaluate QUALIFY before the projection (#29746)
- Type exact SQL numeric literals as Decimal and truncate SQL %/DIV (#29536)
- Allow deterministic expression plugins to opt into CSE/CSPE (#29428)
- Read Parquet ENUM type as pl.String (#29331)
- More map operations (#29296)
- Deprecate cut/qcut (#29329)
- 在聚合行上对分组行运行 SQL 窗口函数,并在投影前评估 QUALIFY(#29746)
- 将精确 SQL 数值字面量键入为 Decimal 并截断 SQL %/DIV(#29536)
- 允许确定性表达式插件选择加入 CSE/CSPE(#29428)
- 将 Parquet ENUM 类型读取为 pl.String(#29331)
- 更多 map 操作(#29296)
- 弃用 cut/qcut(#29329)
🚀 Performance improvements
🚀 性能改进
- Fix phases ending too soon due to join sampling (#29752)
- Group record batch fetches of remote IPC scans (#29750)
- Mark group-by node as memory intensive pipeline blocker (#29749)
- Don't inline slow OOC memory drift path (#29745)
- Only prefetch on unspill when there is memory headroom (#29701)
- Pushdown filters to scan_lance (#29538)
- Do not materialize equal ScalarColumns in Column::append/extend (#28989)
- Evaluate scalar windows in the streaming window node (#29698)
- Don't materialize scalar columns in DataFrame::estimated_size (#29711)
- Defer cached pread from prefetch to decode for Parquet (#29702)
- Decode IPC scans out of order when order is not observed (#29675)
- IPC metadata suffix fetch (#29692)
- Speed up null tracking in streaming group-by sum, min and max (#29691)
- Cache Iceberg manifest files across scans (#29623)
- Improve pushdown support for fallible Hive predicates (#29594)
- Streamline BytecodeParser method rewriting (#29632)
- Stream from_dicts records straight into column buffers (#29656)
- Read correlated SQL aggregates from the outer join when the subquery repeats it (#29685)
- Emit files unordered for remote IPC files when order not observed (#29678)
- Don't keep input order for order-insensitive windows in all engines (#29680)
- Fix parallelization in top-k reducer (#29684)
- Share one bloom filter between the build threads of a runtime filter (#29677)
- Add streaming (out-of-core) sort (#29657)
- Insert keys that start a run of equal keys into the hot group-by table right away (#29673)
- Prefetch key-row hash table lookups in blocks (#29663)
- Lower count-guarded sums to a sum that is null on empty input (#29668)
- Use single PUT for small cloud uploads (#29603)
- Don't cache plain scans in the streaming engine (#29654)
- Change single-key group-by hash table layout and add prefetching (#29661)
- Use the integer ranges of sampled parquet row groups in filter estimates (#29655)
- Raise HTTP rate-limiter read init default (#29584)
- Change join hash table layout and add prefetching (#29653)
- Add a fast-path for as_list with only a single input (#29592)
- Use one runtime filter builder per thread for a finished sample (#29646)
- Compute decimal addition, subtraction and multiplication in i64 when values fit (#29638)
- Spread decimal sums and means over lanes and check the sum's 38 digits once (#29637)
- Compute decimal addition, subtraction and multiplication without per-row validity (#29634)
- Skip or reuse UDF disassembly in BytecodeParser (#29625)
- Rescale a single decimal operand once in addition and subtraction (#29631)
- Dynamic hot table size in streaming GroupBy (#29629)
- Lower gathering from literal as elementwise (#29605)
- Execute fixed-size rolling windows on streaming engine (#29602)
- Use the finite rule of equi joins to determine build side of semi/anti (#29617)
- Drop the byte limit on runtime filter build sides (#29618)
- Add eviction-free path to GroupedReduction, with lanes (#29613)
- Improve window functions (#29591)
- Push down is_nan/is_not_nan filters to Iceberg (#29579)
- Compute shared group-by agg input subexpressions once (#29574)
- Emit files unordered for remote Parquet files when order not observed (#29581)
- Filter rows by the runtime join key range (#29576)
- Update few-group sums, means and counts in lanes (#29573)
- Loop over 4k arrays in group-by update (#29572)
- Decode parquet scans out of order when order is not observed (#29546)
- Don't always row-encode multi-column group-by (#29570)
- Push down str.starts_with filters to Iceberg (#29545)
- Trace join runtime filters through unions (#29565)
- Fold SQL temporal literals into plain values (#29562)
- Don't rehash stored key on every TotalIndexMap probe (#29561)
- Fuse drops into all filters on streaming engine (#29550)
- Join probe on Arrays directly (#29558)
- Better build side selection in equi-join sampler (#29524)
- Large optimisation for BytecodeParser rewrite/dispatch (#29548)
- Parallelize expressions in in-memory map fallback (#29535)
- Push down != filters to Iceberg (#29455)
- Improve predicate stats (#29495)
- Scale bloom filters from keys seen (#29477)
- Optimize simple like queries (#29459)
- Insert head slice below select(len() <cmp> n) to enable slice pushdown (#29209)
- Parallelize row encoding in the row_encode expression (#29434)
- Degrade outer join optimization (#29432)
- Optionally reorder semi join (#29443)
- Borrow the cached regex instead of cloning it per page (#29422)
- Lower SQL LIKE '%x%' to a literal substring match (#29431)
- Optimize is_in expression (#29425)
- Pushdown bloom filters in probe side (#29423)
- Narrow the decimal rescale to 64 bits when the value fits (#29396)
- Use a thread-local copy of a pushed-down regex predicate (#29411)
- Sample anti-join build side (#29416)
- Improve performance importing from arrow (#29412)
- Sample table size to determine build side in semi joins (#29375)
- Optimize row-encoding (#29405)
- Order pushed parquet predicate columns by measured selectivity (#29397)
- Coerce float literals to decimal instead of casting the column (#29395)
- Fix OOM on TPCH SQL and fix fuzzing errors (#29389)
- Restrict correlated SQL aggregates to requested keys (#29383)
- Prepare Parquet scans for row-group splitting in Polars Cloud (#29295)
- Skip row groups by join runtime ranges without a statistics frame (#29370)
- Reduce copy in scan_lines (#29310)
- Increase HTTP read rate-limit default (#29363)
- Evaluate a pushed parquet predicate sequentially per conjuct (#29352)
- Disable system certificates for CloudScheme::Http sources (#29284)
- Reduce rechunk in sort_in_place (#29343)
- Dynamic predicates for hash joins (#29312)
- Use HTTP suffix range for Parquet size and footer (#29308)
- Derive predicates from join conditions (#29304)
- Lower uncorrelated subqueries to semi joins and push semi/anti joins below inner joins (#29289)
- Make leaf name iterator unique (#29291)
- Don't clone the full frame per arm in when/then/otherwise (#29258)
- Push inner joins before outer joins and rewrite left-join-is-null to anti join (#29277)
- Use stats to decide cross join buffering side (#29270)
- Improve cache-removal and join-order cost estimates (#29263)
- Improve CSPE cost evaluation (#29250)
- Fuse group-by pre-select into node after partition (#29251)
- Inline hot small functions (#29244)
- Fix plan-time regressions in projection pushdown for wide frames (#28724)
- Rechunk before selecting in group-by pre-select (#29219)
- Reuse Iceberg data file sizes (#29063)
- 修复因连接采样导致阶段过早结束的问题(#29752)
- 分组获取远程 IPC 扫描的记录批处理(#29750)
- 将 group-by 节点标记为内存密集型管道阻塞器(#29749)
- 不内联缓慢的 OOC 内存漂移路径(#29745)
- 仅在存在内存余量时进行 unspill 预取(#29701)
- 下推过滤器至 scan_lance(#29538)
- 不在 Column::append/extend 中物化相等的 ScalarColumns(#28989)
- 在流式窗口节点中评估标量窗口(#29698)
- 不在 DataFrame::estimated_size 中物化标量列(#29711)
- 推迟 Parquet 的缓存预读从预取到解码(#29702)
- 当不要求顺序时乱序解码 IPC 扫描(#29675)
- IPC 元数据后缀获取(#29692)
- 加速流式 group-by 求和、最小值和最大值中的空值跟踪(#29691)
- 跨扫描缓存 Iceberg manifest 文件(#29623)
- 改进对不可靠 Hive 谓词的下推支持(#29594)
- 简化 BytecodeParser 方法重写(#29632)
- 将 from_dicts 记录直接流式传输到列缓冲区 (#29656)
- 当子查询重复时,从外部连接中读取相关的 SQL 聚合结果 (#29685)
- 当未观察到顺序时,为远程 IPC 文件无序输出文件 (#29678)
- 在所有引擎中,对于顺序不敏感的窗口操作,不保留输入顺序 (#29680)
- 修复 top-k 归约器中的并行化问题 (#29684)
- 在运行时过滤器的构建线程之间共享一个布隆过滤器 (#29677)
- 添加流式(核心外)排序功能 (#29657)
- 立即将启动相等键序列的键插入到热分组表中 (#29673)
- 以块为单位预取键-行哈希表查找 (#29663)
- 将受计数保护的求和转换为空输入时为 null 的普通求和 (#29668)
- 对小规模云上传使用单个 PUT 请求 (#29603)
- 不在流式引擎中缓存普通扫描 (#29654)
- 更改单键分组哈希表的布局并添加预取功能 (#29661)
- 在过滤估计中使用采样 Parquet 行组的整数范围 (#29655)
- 提高 HTTP 速率限制器读取初始化的默认值 (#29584)
- 更改连接哈希表的布局并添加预取功能 (#29653)
- 为仅有一个输入的 as_list 添加快速路径 (#29592)
- 为已完成的样本,每个线程使用一个运行时过滤器构建器 (#29646)
- 当数值适配时,使用 i64 计算十进制加法、减法和乘法 (#29638)
- 将十进制求和与均值分散到多个通道,并仅检查一次求和结果的 38 位数字 (#29637)
- 在不逐行检查有效性的情况下计算十进制加法、减法和乘法 (#29634)
- 在 BytecodeParser 中跳过或重用 UDF 反汇编 (#29625)
- 在加法和减法中仅对单个十进制操作数重新缩放一次 (#29631)
- 流式 GroupBy 中的动态热表大小 (#29629)
- 从字面量作为元素执行降低收集 (#29605)
- 在流式引擎上执行固定大小的滚动窗口 (#29602)
- 使用等值连接的限制规则来确定半连接/反连接的构建侧 (#29617)
- 移除运行时过滤器构建侧的字节限制 (#29618)
- 为 GroupedReduction 添加无驱逐路径,带有通道 (lanes) (#29613)
- 改进窗口函数 (#29591)
- 将 is_nan/is_not_nan 过滤器下推到 Iceberg (#29579)
- 仅计算一次共享 group-by 聚合输入子表达式 (#29574)
- 当未观察到顺序时,无序输出远程 Parquet 文件 (#29581)
- 通过运行时连接键范围过滤行 (#29576)
- 在通道 (lanes) 中更新少量组的和、均值和计数 (#29573)
- 在 group-by 更新中遍历 4k 数组 (#29572)
- 当未观察到顺序时,无序解码 Parquet 扫描结果 (#29546)
- 不要总是对多列 group-by 进行行编码 (#29570)
- 将 str.starts_with 过滤器下推到 Iceberg (#29545)
- 通过联合 (unions) 追踪连接运行时过滤器 (#29565)
- 将 SQL 时间字面量折叠为普通值 (#29562)
- 不要在每次 TotalIndexMap 探测时重新哈希存储的键 (#29561)
- 在流式引擎上将丢弃操作融合到所有过滤器中 (#29550)
- 直接在 Arrays 上进行连接探测 (#29558)
- 在等值连接采样器中更好地选择构建侧 (#29524)
- BytecodeParser 重写/分发的重大优化 (#29548)
- 在内存映射回退中并行化表达式 (#29535)
- 将 != 过滤器下推到 Iceberg (#29455)
- 改进谓词统计信息 (#29495)
- 根据已看到的键扩展布隆过滤器 (#29477)
- 优化简单的 LIKE 查询 (#29459)
- 在 select(len() <cmp> n) 下方插入头部切片以启用切片下推 (#29209)
- 并行化 row_encode 表达式中的行编码 (#29434)
- 降级外连接优化 (#29432)
- 可选地重排序半连接 (#29443)
- 借用缓存的正则表达式,而不是每页克隆一次 (#29422)
- 将 SQL LIKE '%x%' 降级为字面量子串匹配 (#29431)
- 优化 is_in 表达式 (#29425)
- 在下推的探测侧应用布隆过滤器 (#29423)
- 当值适合时,将小数重新缩放范围缩小至 64 位 (#29396)
- 使用下推正则表达式谓词的线程本地副本 (#29411)
- 采样反连接的构建侧 (#29416)
- 提高从 Arrow 导入的性能 (#29412)
- 采样表大小以确定半连接中的构建侧 (#29375)
- 优化行编码 (#29405)
- 按测量的选择性对下推的 Parquet 谓词列进行排序 (#29397)
- 将浮点字面量强制转换为小数,而不是转换列 (#29395)
- 修复 TPCH SQL 上的内存溢出 (OOM) 并修复模糊测试错误 (#29389)
- 将相关 SQL 聚合限制为请求的键 (#29383)
- 为 Polars Cloud 中的行组拆分准备 Parquet 扫描 (#29295)
- 在没有统计信息帧的情况下通过连接运行时范围跳过行组 (#29370)
- 减少 scan_lines 中的拷贝操作 (#29310)
- 提高 HTTP 读取速率限制默认值 (#29363)
- 按合取项顺序逐个评估推送的 Parquet 谓词 (#29352)
- 为 CloudScheme::Http 源禁用系统证书 (#29284)
- 减少 sort_in_place 中的重新分块操作 (#29343)
- 哈希连接中的动态谓词 (#29312)
- 对 Parquet 大小和页尾使用 HTTP 后缀范围 (#29308)
- 从连接条件推导谓词 (#29304)
- 将非相关子查询下推为半连接,并将半连接/反连接下推到内连接之下 (#29289)
- 使叶子名称迭代器唯一化 (#29291)
- 在 when/then/otherwise 中不要为每个分支克隆完整的帧 (#29258)
- 先于外连接处理内连接,并将左连接为空重写为反连接 (#29277)
- 使用统计信息决定交叉连接的缓冲侧 (#29270)
- 改进缓存移除和连接顺序的成本估算 (#29263)
- 改进 CSPE 成本评估 (#29250)
- 在分区后的节点中将分组预选择融合到分组操作中 (#29251)
- 内联热点小函数 (#29244)
- 修复宽帧投影下推中的计划时回归问题 (#28724)
- 在分组预选择的选取之前进行重新分块 (#29219)
- 复用 Iceberg 数据文件大小 (#29063)
✨ Enhancements
✨ 增强功能
- Python Polars 2.0 (#29723)
- Enable OOC by default with 80% of available RAM as threshold (#29741)
- Support window frames in SQL FIRST_VALUE and LAST_VALUE, and add NTH_VALUE and a LAG/LEAD default (#29740)
- Set default OOC disk budget to 64 GB (#29734)
- Add POLARS_OOMKILL_THRESHOLD_MB (#29735)
- Improve estimated DataFrame memory usage (#29712)
- Extend pipe_with_dtype for multiple expressions (#29676)
- Attribute physical nodes to IR nodes (#29522)
- Make approx_quantile sketch states mergeable across processes (#29660)
- Add unstable scan_lance (#29413)
- New Expression and function: pipe_with_dtype (#29547)
- Resolve is_in and Map lookup coercion from dtypes alone (#29487)
- Cast the needle of is_in and Map lookups exactly or not at all (#29486)
- Support Map in JSON and NDJSON reading and writing (#29478)
- Add scan_external_reader (#29232)
- Python Polars 2.0 (#29723)
- 默认启用 OOC(Out-of-Core),以可用 RAM 的 80% 作为阈值 (#29741)
- 支持 SQL FIRST_VALUE 和 LAST_VALUE 中的窗口帧,并添加 NTH_VALUE 以及 LAG/LEAD 的默认值 (#29740)
- 将默认 OOC 磁盘预算设置为 64 GB (#29734)
- 添加 POLARS_OOMKILL_THRESHOLD_MB (#29735)
- 改进 DataFrame 内存使用量的估算 (#29712)
- 扩展 pipe_with_dtype 以支持多个表达式 (#29676)
- 将物理节点映射到 IR 节点 (#29522)
- 使 approx_quantile 草图状态在进程间可合并 (#29660)
- 添加不稳定的 scan_lance (#29413)
- 新增 Expression 和函数:pipe_with_dtype (#29547)
- 仅从数据类型解析 is_in 和 Map 查找的强制转换 (#29487)
- 对 is_in 和 Map 查找中的 needle 进行精确或完全不进行类型转换 (#29486)
- 在 JSON 和 NDJSON 的读写中支持 Map (#29478)
- 添加 scan_external_reader (#29232)
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力