复杂文字版面还原与多语种文档转写重构
Tesseract OCR
Google Cloud Vision API
python-docx
Unicode Normalization
Microsoft Word
正在基于该需求推演落地方案…
需求痛点(公开)
通用 OCR 与 PDF 转换工具在处理印地语等复杂文字时,极易产生连字识别错误与符号乱码,且无法保留书籍原本的标题层级、字体样式和分页结构,导致人工重新录入排版的成本极为高昂。
切入方向(公开)
结合语种专属 OCR 引擎与 Unicode 字符规范化方案,配合版面分析与样式映射算法,将扫描书籍高保真重构为原生 Microsoft Word(.docx)格式,完整复刻原书排版与层级样式。
原始需求溯源 · 2 条
"I have the entire Hindi book in a PDF file and I need every word faithfully re-entered into a Word document. The goal is a clean, plain-text file—no…"
"I have a complete Hindi book in PDF format. The scan is crystal-clear and every page is readable, so you will be able to see the original layout in…"
仅展示任务摘要与外链,不转载原文全文;个人信息已脱敏。数据来源已登记,可溯源。
💬 社区讨论