FineBooks项目测试14款OCR模型,提升历史书籍文本质量
Old OCR text cripples language model training, and FineBooks wants to fix that at scale
The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says.
来自Hugging Face和EleutherAI的FineBooks项目在超过2000页历史书籍上测试了14个开源OCR模型。表现最好的模型dots.mocr在每千页成本不到两美元的情况下,字符准确率达到97.6%。团队表示,这对于AI训练数据来说已经足够好,但对于学术转录来说还不够。
The article Old OCR text cripples language model training, and FineBooks wants to fix that at scale appeared first on The Decoder.
文章《旧OCR文本阻碍语言模型训练,FineBooks希望大规模解决这一问题》首次出现在The Decoder上。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力