精选80Rohan Paul模型发布/更新
Kimi K3 生成 H100 CUDA 内核速度比优化 PyTorch 快 14.82 倍
Kimi K3 now writes a H100 CUDA kernel 14.82x faster than optimized PyTorch, near…
Kimi K3 now writes a H100 CUDA kernel 14.82x faster than optimized PyTorch, nearly matching Claude Opus 4.8.
KernelBench tests hardware work, not just whether generated code successfully compiles.
KernelBench gives an AI agent a readable PyTorch implementation and asks it to produce a custom GPU implementation that is both correct and faster.
The code must compile, return the right numerical results across test inputs and beat a reference implementation on actual hardware. A kernel that is fast but wrong fails. A kernel that is correct but slow is not particularly useful.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力