精选75jietang论文研究
SAO异步优化算法在编码推理基准上超越GRPO
Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for…
Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. https://arxiv.org/pdf/2607.07508
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力