精选90OpenAI News(RSS)模型发布/更新
OpenAI 发布 PPO 强化学习算法
推荐理由
PPO 是强化学习领域的里程碑式算法,做 RL 的同学必须了解,建议直接阅读原文并尝试复现。
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力