跳到主内容
@wquguru
精选90Polars(GitHub Releases)语言与库

Python Polars v2.0.0:破坏性变更与弃用 cut/qcut

Python Polars 2.0.0

原文
发到 X
推荐理由

v2.0.0 包含多项破坏性变更(如 SQL 窗口函数行为改变、ENUM 类型解析变化)及核心 API 弃用,直接打断现有用法。请检查代码中是否依赖旧版 SQL 语义或使用了已弃用的 cut/qcut,并尽快适配。

💥 Breaking changes

💥 破坏性变更

  • Run SQL window functions over grouped rows on the aggregated rows, and evaluate QUALIFY before the projection (#29746)
  • Type exact SQL numeric literals as Decimal and truncate SQL %/DIV (#29536)
  • Allow deterministic expression plugins to opt into CSE/CSPE (#29428)
  • Read Parquet ENUM type as pl.String (#29331)
  • More map operations (#29296)
  • Deprecate cut/qcut (#29329)
  • 在聚合行上对分组行运行 SQL 窗口函数,并在投影前评估 QUALIFY(#29746)
  • 将精确 SQL 数值字面量键入为 Decimal 并截断 SQL %/DIV(#29536)
  • 允许确定性表达式插件选择加入 CSE/CSPE(#29428)
  • 将 Parquet ENUM 类型读取为 pl.String(#29331)
  • 更多 map 操作(#29296)
  • 弃用 cut/qcut(#29329)

🚀 Performance improvements

🚀 性能改进

  • Fix phases ending too soon due to join sampling (#29752)
  • Group record batch fetches of remote IPC scans (#29750)
  • Mark group-by node as memory intensive pipeline blocker (#29749)
  • Don't inline slow OOC memory drift path (#29745)
  • Only prefetch on unspill when there is memory headroom (#29701)
  • Pushdown filters to scan_lance (#29538)
  • Do not materialize equal ScalarColumns in Column::append/extend (#28989)
  • Evaluate scalar windows in the streaming window node (#29698)
  • Don't materialize scalar columns in DataFrame::estimated_size (#29711)
  • Defer cached pread from prefetch to decode for Parquet (#29702)
  • Decode IPC scans out of order when order is not observed (#29675)
  • IPC metadata suffix fetch (#29692)
  • Speed up null tracking in streaming group-by sum, min and max (#29691)
  • Cache Iceberg manifest files across scans (#29623)
  • Improve pushdown support for fallible Hive predicates (#29594)
  • Streamline BytecodeParser method rewriting (#29632)
  • Stream from_dicts records straight into column buffers (#29656)
  • Read correlated SQL aggregates from the outer join when the subquery repeats it (#29685)
  • Emit files unordered for remote IPC files when order not observed (#29678)
  • Don't keep input order for order-insensitive windows in all engines (#29680)
  • Fix parallelization in top-k reducer (#29684)
  • Share one bloom filter between the build threads of a runtime filter (#29677)
  • Add streaming (out-of-core) sort (#29657)
  • Insert keys that start a run of equal keys into the hot group-by table right away (#29673)
  • Prefetch key-row hash table lookups in blocks (#29663)
  • Lower count-guarded sums to a sum that is null on empty input (#29668)
  • Use single PUT for small cloud uploads (#29603)
  • Don't cache plain scans in the streaming engine (#29654)
  • Change single-key group-by hash table layout and add prefetching (#29661)
  • Use the integer ranges of sampled parquet row groups in filter estimates (#29655)
  • Raise HTTP rate-limiter read init default (#29584)
  • Change join hash table layout and add prefetching (#29653)
  • Add a fast-path for as_list with only a single input (#29592)
  • Use one runtime filter builder per thread for a finished sample (#29646)
  • Compute decimal addition, subtraction and multiplication in i64 when values fit (#29638)
  • Spread decimal sums and means over lanes and check the sum's 38 digits once (#29637)
  • Compute decimal addition, subtraction and multiplication without per-row validity (#29634)
  • Skip or reuse UDF disassembly in BytecodeParser (#29625)
  • Rescale a single decimal operand once in addition and subtraction (#29631)
  • Dynamic hot table size in streaming GroupBy (#29629)
  • Lower gathering from literal as elementwise (#29605)
  • Execute fixed-size rolling windows on streaming engine (#29602)
  • Use the finite rule of equi joins to determine build side of semi/anti (#29617)
  • Drop the byte limit on runtime filter build sides (#29618)
  • Add eviction-free path to GroupedReduction, with lanes (#29613)
  • Improve window functions (#29591)
  • Push down is_nan/is_not_nan filters to Iceberg (#29579)
  • Compute shared group-by agg input subexpressions once (#29574)
  • Emit files unordered for remote Parquet files when order not observed (#29581)
  • Filter rows by the runtime join key range (#29576)
  • Update few-group sums, means and counts in lanes (#29573)
  • Loop over 4k arrays in group-by update (#29572)
  • Decode parquet scans out of order when order is not observed (#29546)
  • Don't always row-encode multi-column group-by (#29570)
  • Push down str.starts_with filters to Iceberg (#29545)
  • Trace join runtime filters through unions (#29565)
  • Fold SQL temporal literals into plain values (#29562)
  • Don't rehash stored key on every TotalIndexMap probe (#29561)
  • Fuse drops into all filters on streaming engine (#29550)
  • Join probe on Arrays directly (#29558)
  • Better build side selection in equi-join sampler (#29524)
  • Large optimisation for BytecodeParser rewrite/dispatch (#29548)
  • Parallelize expressions in in-memory map fallback (#29535)
  • Push down != filters to Iceberg (#29455)
  • Improve predicate stats (#29495)
  • Scale bloom filters from keys seen (#29477)
  • Optimize simple like queries (#29459)
  • Insert head slice below select(len() <cmp> n) to enable slice pushdown (#29209)
  • Parallelize row encoding in the row_encode expression (#29434)
  • Degrade outer join optimization (#29432)
  • Optionally reorder semi join (#29443)
  • Borrow the cached regex instead of cloning it per page (#29422)
  • Lower SQL LIKE '%x%' to a literal substring match (#29431)
  • Optimize is_in expression (#29425)
  • Pushdown bloom filters in probe side (#29423)
  • Narrow the decimal rescale to 64 bits when the value fits (#29396)
  • Use a thread-local copy of a pushed-down regex predicate (#29411)
  • Sample anti-join build side (#29416)
  • Improve performance importing from arrow (#29412)
  • Sample table size to determine build side in semi joins (#29375)
  • Optimize row-encoding (#29405)
  • Order pushed parquet predicate columns by measured selectivity (#29397)
  • Coerce float literals to decimal instead of casting the column (#29395)
  • Fix OOM on TPCH SQL and fix fuzzing errors (#29389)
  • Restrict correlated SQL aggregates to requested keys (#29383)
  • Prepare Parquet scans for row-group splitting in Polars Cloud (#29295)
  • Skip row groups by join runtime ranges without a statistics frame (#29370)
  • Reduce copy in scan_lines (#29310)
  • Increase HTTP read rate-limit default (#29363)
  • Evaluate a pushed parquet predicate sequentially per conjuct (#29352)
  • Disable system certificates for CloudScheme::Http sources (#29284)
  • Reduce rechunk in sort_in_place (#29343)
  • Dynamic predicates for hash joins (#29312)
  • Use HTTP suffix range for Parquet size and footer (#29308)
  • Derive predicates from join conditions (#29304)
  • Lower uncorrelated subqueries to semi joins and push semi/anti joins below inner joins (#29289)
  • Make leaf name iterator unique (#29291)
  • Don't clone the full frame per arm in when/then/otherwise (#29258)
  • Push inner joins before outer joins and rewrite left-join-is-null to anti join (#29277)
  • Use stats to decide cross join buffering side (#29270)
  • Improve cache-removal and join-order cost estimates (#29263)
  • Improve CSPE cost evaluation (#29250)
  • Fuse group-by pre-select into node after partition (#29251)
  • Inline hot small functions (#29244)
  • Fix plan-time regressions in projection pushdown for wide frames (#28724)
  • Rechunk before selecting in group-by pre-select (#29219)
  • Reuse Iceberg data file sizes (#29063)
  • 修复因连接采样导致阶段过早结束的问题(#29752)
  • 分组获取远程 IPC 扫描的记录批处理(#29750)
  • 将 group-by 节点标记为内存密集型管道阻塞器(#29749)
  • 不内联缓慢的 OOC 内存漂移路径(#29745)
  • 仅在存在内存余量时进行 unspill 预取(#29701)
  • 下推过滤器至 scan_lance(#29538)
  • 不在 Column::append/extend 中物化相等的 ScalarColumns(#28989)
  • 在流式窗口节点中评估标量窗口(#29698)
  • 不在 DataFrame::estimated_size 中物化标量列(#29711)
  • 推迟 Parquet 的缓存预读从预取到解码(#29702)
  • 当不要求顺序时乱序解码 IPC 扫描(#29675)
  • IPC 元数据后缀获取(#29692)
  • 加速流式 group-by 求和、最小值和最大值中的空值跟踪(#29691)
  • 跨扫描缓存 Iceberg manifest 文件(#29623)
  • 改进对不可靠 Hive 谓词的下推支持(#29594)
  • 简化 BytecodeParser 方法重写(#29632)
  • 将 from_dicts 记录直接流式传输到列缓冲区 (#29656)
  • 当子查询重复时,从外部连接中读取相关的 SQL 聚合结果 (#29685)
  • 当未观察到顺序时,为远程 IPC 文件无序输出文件 (#29678)
  • 在所有引擎中,对于顺序不敏感的窗口操作,不保留输入顺序 (#29680)
  • 修复 top-k 归约器中的并行化问题 (#29684)
  • 在运行时过滤器的构建线程之间共享一个布隆过滤器 (#29677)
  • 添加流式(核心外)排序功能 (#29657)
  • 立即将启动相等键序列的键插入到热分组表中 (#29673)
  • 以块为单位预取键-行哈希表查找 (#29663)
  • 将受计数保护的求和转换为空输入时为 null 的普通求和 (#29668)
  • 对小规模云上传使用单个 PUT 请求 (#29603)
  • 不在流式引擎中缓存普通扫描 (#29654)
  • 更改单键分组哈希表的布局并添加预取功能 (#29661)
  • 在过滤估计中使用采样 Parquet 行组的整数范围 (#29655)
  • 提高 HTTP 速率限制器读取初始化的默认值 (#29584)
  • 更改连接哈希表的布局并添加预取功能 (#29653)
  • 为仅有一个输入的 as_list 添加快速路径 (#29592)
  • 为已完成的样本,每个线程使用一个运行时过滤器构建器 (#29646)
  • 当数值适配时,使用 i64 计算十进制加法、减法和乘法 (#29638)
  • 将十进制求和与均值分散到多个通道,并仅检查一次求和结果的 38 位数字 (#29637)
  • 在不逐行检查有效性的情况下计算十进制加法、减法和乘法 (#29634)
  • 在 BytecodeParser 中跳过或重用 UDF 反汇编 (#29625)
  • 在加法和减法中仅对单个十进制操作数重新缩放一次 (#29631)
  • 流式 GroupBy 中的动态热表大小 (#29629)
  • 从字面量作为元素执行降低收集 (#29605)
  • 在流式引擎上执行固定大小的滚动窗口 (#29602)
  • 使用等值连接的限制规则来确定半连接/反连接的构建侧 (#29617)
  • 移除运行时过滤器构建侧的字节限制 (#29618)
  • 为 GroupedReduction 添加无驱逐路径,带有通道 (lanes) (#29613)
  • 改进窗口函数 (#29591)
  • 将 is_nan/is_not_nan 过滤器下推到 Iceberg (#29579)
  • 仅计算一次共享 group-by 聚合输入子表达式 (#29574)
  • 当未观察到顺序时,无序输出远程 Parquet 文件 (#29581)
  • 通过运行时连接键范围过滤行 (#29576)
  • 在通道 (lanes) 中更新少量组的和、均值和计数 (#29573)
  • 在 group-by 更新中遍历 4k 数组 (#29572)
  • 当未观察到顺序时,无序解码 Parquet 扫描结果 (#29546)
  • 不要总是对多列 group-by 进行行编码 (#29570)
  • 将 str.starts_with 过滤器下推到 Iceberg (#29545)
  • 通过联合 (unions) 追踪连接运行时过滤器 (#29565)
  • 将 SQL 时间字面量折叠为普通值 (#29562)
  • 不要在每次 TotalIndexMap 探测时重新哈希存储的键 (#29561)
  • 在流式引擎上将丢弃操作融合到所有过滤器中 (#29550)
  • 直接在 Arrays 上进行连接探测 (#29558)
  • 在等值连接采样器中更好地选择构建侧 (#29524)
  • BytecodeParser 重写/分发的重大优化 (#29548)
  • 在内存映射回退中并行化表达式 (#29535)
  • 将 != 过滤器下推到 Iceberg (#29455)
  • 改进谓词统计信息 (#29495)
  • 根据已看到的键扩展布隆过滤器 (#29477)
  • 优化简单的 LIKE 查询 (#29459)
  • 在 select(len() <cmp> n) 下方插入头部切片以启用切片下推 (#29209)
  • 并行化 row_encode 表达式中的行编码 (#29434)
  • 降级外连接优化 (#29432)
  • 可选地重排序半连接 (#29443)
  • 借用缓存的正则表达式,而不是每页克隆一次 (#29422)
  • 将 SQL LIKE '%x%' 降级为字面量子串匹配 (#29431)
  • 优化 is_in 表达式 (#29425)
  • 在下推的探测侧应用布隆过滤器 (#29423)
  • 当值适合时,将小数重新缩放范围缩小至 64 位 (#29396)
  • 使用下推正则表达式谓词的线程本地副本 (#29411)
  • 采样反连接的构建侧 (#29416)
  • 提高从 Arrow 导入的性能 (#29412)
  • 采样表大小以确定半连接中的构建侧 (#29375)
  • 优化行编码 (#29405)
  • 按测量的选择性对下推的 Parquet 谓词列进行排序 (#29397)
  • 将浮点字面量强制转换为小数,而不是转换列 (#29395)
  • 修复 TPCH SQL 上的内存溢出 (OOM) 并修复模糊测试错误 (#29389)
  • 将相关 SQL 聚合限制为请求的键 (#29383)
  • 为 Polars Cloud 中的行组拆分准备 Parquet 扫描 (#29295)
  • 在没有统计信息帧的情况下通过连接运行时范围跳过行组 (#29370)
  • 减少 scan_lines 中的拷贝操作 (#29310)
  • 提高 HTTP 读取速率限制默认值 (#29363)
  • 按合取项顺序逐个评估推送的 Parquet 谓词 (#29352)
  • 为 CloudScheme::Http 源禁用系统证书 (#29284)
  • 减少 sort_in_place 中的重新分块操作 (#29343)
  • 哈希连接中的动态谓词 (#29312)
  • 对 Parquet 大小和页尾使用 HTTP 后缀范围 (#29308)
  • 从连接条件推导谓词 (#29304)
  • 将非相关子查询下推为半连接,并将半连接/反连接下推到内连接之下 (#29289)
  • 使叶子名称迭代器唯一化 (#29291)
  • 在 when/then/otherwise 中不要为每个分支克隆完整的帧 (#29258)
  • 先于外连接处理内连接,并将左连接为空重写为反连接 (#29277)
  • 使用统计信息决定交叉连接的缓冲侧 (#29270)
  • 改进缓存移除和连接顺序的成本估算 (#29263)
  • 改进 CSPE 成本评估 (#29250)
  • 在分区后的节点中将分组预选择融合到分组操作中 (#29251)
  • 内联热点小函数 (#29244)
  • 修复宽帧投影下推中的计划时回归问题 (#28724)
  • 在分组预选择的选取之前进行重新分块 (#29219)
  • 复用 Iceberg 数据文件大小 (#29063)

✨ Enhancements

✨ 增强功能

  • Python Polars 2.0 (#29723)
  • Enable OOC by default with 80% of available RAM as threshold (#29741)
  • Support window frames in SQL FIRST_VALUE and LAST_VALUE, and add NTH_VALUE and a LAG/LEAD default (#29740)
  • Set default OOC disk budget to 64 GB (#29734)
  • Add POLARS_OOMKILL_THRESHOLD_MB (#29735)
  • Improve estimated DataFrame memory usage (#29712)
  • Extend pipe_with_dtype for multiple expressions (#29676)
  • Attribute physical nodes to IR nodes (#29522)
  • Make approx_quantile sketch states mergeable across processes (#29660)
  • Add unstable scan_lance (#29413)
  • New Expression and function: pipe_with_dtype (#29547)
  • Resolve is_in and Map lookup coercion from dtypes alone (#29487)
  • Cast the needle of is_in and Map lookups exactly or not at all (#29486)
  • Support Map in JSON and NDJSON reading and writing (#29478)
  • Add scan_external_reader (#29232)
  • Python Polars 2.0 (#29723)
  • 默认启用 OOC(Out-of-Core),以可用 RAM 的 80% 作为阈值 (#29741)
  • 支持 SQL FIRST_VALUE 和 LAST_VALUE 中的窗口帧,并添加 NTH_VALUE 以及 LAG/LEAD 的默认值 (#29740)
  • 将默认 OOC 磁盘预算设置为 64 GB (#29734)
  • 添加 POLARS_OOMKILL_THRESHOLD_MB (#29735)
  • 改进 DataFrame 内存使用量的估算 (#29712)
  • 扩展 pipe_with_dtype 以支持多个表达式 (#29676)
  • 将物理节点映射到 IR 节点 (#29522)
  • 使 approx_quantile 草图状态在进程间可合并 (#29660)
  • 添加不稳定的 scan_lance (#29413)
  • 新增 Expression 和函数:pipe_with_dtype (#29547)
  • 仅从数据类型解析 is_in 和 Map 查找的强制转换 (#29487)
  • 对 is_in 和 Map 查找中的 needle 进行精确或完全不进行类型转换 (#29486)
  • 在 JSON 和 NDJSON 的读写中支持 Map (#29478)
  • 添加 scan_external_reader (#29232)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件