现代NLP分词综述:32位研究者历时8个月完成
Tokenization: A Survey for Modern NLP [R]
分词是大模型底层关键但常被忽视的环节,这份由32位专家耗时8个月完成的全面综述极具参考价值,建议LLM研发人员重点阅读以理解技术演进与选型权衡。
Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field.
尽管标记化(Tokenization)对自然语言处理(NLP)的各个方面都有影响,但它仍是语言建模中一个被严重低估的研究领域。在过去约8个月的时间里,32位(!)标记化研究人员共同完成了该领域最全面的综述。
We cover every aspect of tokenization: algorithms, evaluations, multilinguality, encodings, theory, etc. We even cover what you might want to replace tokenizers with (e.g., latent or visual tokenization). We also cover some topics that are closely adjacent to tokenization, such as constrained generation, token healing, and tokenizer security concerns.
我们涵盖了标记化的所有方面:算法、评估、多语言能力、编码、理论等。我们甚至探讨了你可能用来替代标记化的方法(例如潜在标记化或视觉标记化)。我们还涵盖了一些与标记化紧密相关的主题,如约束生成、标记修复以及标记化器的安全问题。
Check it out!
快来看看吧!
https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力