Perceptron发布Mk1.5具身智能模型,首人称视频理解能力跃升
Perceptron released Mk1.5, a 35B embodied-agent model with up to 4.7× faster req…
Mk1.5展示了通过专用数据训练和工具调用实现的能力跃迁,为具身智能提供了可复用的工程思路,值得关注。
Perceptron released Mk1.5, a 35B embodied-agent model with up to 4.7× faster request completion.
Perceptron 发布了 Mk1.5,这是一个拥有高达 4.7 倍更快请求完成速度的 35B 具身智能体模型。
4. First-person video is a very different problem from ordinary video understanding.
4. 第一人称视频与普通视频理解是截然不同的问题。
For a robot or a pair of smart glasses, hands constantly enter the frame, manipulate objects, disappear, return, and partially occlude what matters. The camera itself is moving too.
对于机器人或一副智能眼镜来说,双手不断进入画面,操作物体,消失又出现,并部分遮挡关键内容。摄像机本身也在移动。
Perceptron trained Mk1.5 specifically on egocentric video, and very interestingly the model gets a 36.1-point jump on MMSearch when tools are enabled. Moves from 18.2 to 54.3 F1 when it can use web search, page reading, reverse-image search, and zoom.
Perceptron 专门针对自我中心视角视频训练了 Mk1.5,有趣的是,当启用工具时,该模型在 MMSearch 上的得分提升了 36.1 分。在使用网页搜索、页面阅读、反向图像搜索和缩放功能时,其 F1 分数从 18.2 提升至 54.3。
This is a useful way to think about embodied intelligence too.
这也是思考具身智能的一种有用方式。
A robot does not need every possible capability encoded into its weights. It needs to recognize when its current information is insufficient, pick the right external capability, use it, then continue the task.
机器人不需要将所有可能的能力都编码到其权重中。它需要能够识别当前信息是否不足,选择合适的外部能力,加以利用,然后继续执行任务。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力