技术路线新突破:阿里云开源发布0.8B文档解析模型OvisOCR2

凤凰网科技
Jul 24

凤凰网科技讯 7月24日,阿里云正式开源发布文档解析模型OvisOCR2。该模型以综合得分96.58刷新OmniDocBench v1.6榜单纪录,成为首个超越传统流水线方法并登顶该榜单的端到端模型,标志着文档解析领域迎来重要技术路线突破。

此前文档解析领域由“流水线方法”主导,需依次调用版面分析、文字识别等多个模型,存在维护成本高、误差累积等问题。OvisOCR2采用端到端路线:输入文档图像后,模型一次性输出符合自然阅读顺序的Markdown格式结果,覆盖文本、公式、表格和视觉区域,一个模型完成过去多模型协作的工作。

该模型由Qwen3.5-0.8B后训练得到,仅0.8B参数量即实现SOTA性能。训练采用监督微调、强化学习、在线策略蒸馏和模型融合四阶段流程,并构建了真实与合成文档互补的数据引擎体系,以充分释放紧凑模型潜力。OvisOCR2已以Apache 2.0许可证开源,兼容vLLM等主流推理框架,可与Qwen3.5生态无缝衔接。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10