夜雨聆风学习资料网

ARTICLE · 1153156

用 Python 把 PDF 批量转成 Word/Excel,100 份文件 3 分钟搞定

用 Python 把 PDF 批量转成 Word/Excel,100 份文件 3 分钟搞定

打工人自救指南:不再手动复制粘贴

一、先说痛点

上周同事发来一个压缩包,里面 87 份 PDF 合同,要求"整理成 Excel 表格"。

手动操作的话,大概是这样的流程:打开 PDF → 复制 → 粘贴到 Word → 调整格式 → 再复制到 Excel → 下一份……

按每份 3 分钟算,87 份就是 4 个多小时。而且复制过去的表格经常错位,改格式比重新录入还累。

其实这类活儿,Python 几行代码就能解决。今天分享两个库,分别搞定 PDF→Word 和 PDF→Excel。

二、环境准备

pip install pdf2docx pdfplumber pandas openpyxl

四个库各司其职:

库作用
pdf2docxPDF 转 Word,保留排版
pdfplumber提取 PDF 中的表格
pandas数据处理 + 写 Excel
openpyxlExcel 写入引擎

三、PDF 转 Word:pdf2docx

这个库是我用过转换效果最好的,基于 PyMuPDF,能保留文字、图片、表格的大致版式。

from pdf2docx import Converterdef pdf_to_word(pdf_path, docx_path):    cv = Converter(pdf_path)    cv.convert(docx_path, start=0, end=None)  # end=None 表示全部页    cv.close()    print(f"转换完成:{docx_path}")if __name__ == "__main__":    pdf_to_word("合同.pdf", "合同.docx")

三行核心代码,就这一步。

convert() 还支持指定页码范围,比如只转前 5 页:

cv.convert("out.docx", start=0, end=5)

批量转换整个文件夹

from pathlib import Pathfrom pdf2docx import Converterdef batch_pdf_to_word(src_dir, out_dir):    src_dir, out_dir = Path(src_dir), Path(out_dir)    out_dir.mkdir(parents=True, exist_ok=True)    pdf_files = list(src_dir.glob("*.pdf"))    print(f"共发现 {len(pdf_files)} 个 PDF 文件")    for i, pdf_file in enumerate(pdf_files, 1):        docx_file = out_dir / f"{pdf_file.stem}.docx"        try:            cv = Converter(str(pdf_file))            cv.convert(str(docx_file))            cv.close()            print(f"[{i}/{len(pdf_files)}] ✓ {pdf_file.name}")        except Exception as e:            print(f"[{i}/{len(pdf_files)}] ✗ {pdf_file.name} -> {e}")batch_pdf_to_word("./pdfs", "./docs")

87 份文件实测大概 3 分钟跑完,期间你可以去泡杯咖啡。

四、PDF 转 Excel:pdfplumber

Word 解决了,但提取表格数据其实用 Excel 更合适。这里用 pdfplumber,纯 Python 实现,不依赖 Java 环境(比 tabula-py 省事)。

import pdfplumberimport pandas as pddef pdf_to_excel(pdf_path, excel_path):    with pdfplumber.open(pdf_path) as pdf:        with pd.ExcelWriter(excel_path, engine="openpyxl") as writer:            for page_no, page in enumerate(pdf.pages, 1):                tables = page.extract_tables()                for t_no, table in enumerate(tables, 1):                    if not table or len(table) < 2:                        continue                    # 第一行当表头,其余当数据                    header = [str(c).strip() if c else f"列{i+1}"                              for i, c in enumerate(table[0])]                    df = pd.DataFrame(table[1:], columns=header)                    # 清理换行符                    df = df.applymap(                        lambda x: str(x).replace("\n", " ").strip() if x else ""                    )                    sheet_name = f"P{page_no}_T{t_no}"                    df.to_excel(writer, sheet_name=sheet_name, index=False)                    print(f"  第 {page_no} 页 表格{t_no}:{len(df)} 行")    print(f"完成:{excel_path}")pdf_to_excel("报表.pdf", "报表.xlsx")

每个表格写入一个独立的 sheet,命名规则是 P页码_T表序号,方便回溯。

五、合并成一个命令行工具

把两个功能封装起来,支持按后缀自动判断:

import argparsefrom pathlib import Pathfrom pdf2docx import Converterimport pdfplumberimport pandas as pddef pdf2word(pdf_path, out_path):    cv = Converter(str(pdf_path))    cv.convert(str(out_path))    cv.close()def pdf2excel(pdf_path, out_path):    with pdfplumber.open(pdf_path) as pdf:        with pd.ExcelWriter(out_path, engine="openpyxl") as writer:            for p_no, page in enumerate(pdf.pages, 1):                for t_no, table in enumerate(page.extract_tables(), 1):                    if not table or len(table) < 2:                        continue                    header = [str(c).strip() if c else f"列{i+1}"                              for i, c in enumerate(table[0])]                    df = pd.DataFrame(table[1:], columns=header)                    df.to_excel(writer, sheet_name=f"P{p_no}_T{t_no}",                                index=False)def main():    parser = argparse.ArgumentParser(description="PDF 批量转换工具")    parser.add_argument("src", help="PDF 所在文件夹")    parser.add_argument("-o", "--out", default="./output", help="输出目录")    parser.add_argument("-m", "--mode", choices=["word", "excel"],                        default="word", help="转换目标格式")    args = parser.parse_args()    src_dir, out_dir = Path(args.src), Path(args.out)    out_dir.mkdir(parents=True, exist_ok=True)    pdfs = sorted(src_dir.glob("*.pdf"))    print(f"共 {len(pdfs)} 个文件,目标格式:{args.mode}\n")    for i, pdf in enumerate(pdfs, 1):        suffix = ".docx" if args.mode == "word" else ".xlsx"        out_file = out_dir / f"{pdf.stem}{suffix}"        try:            if args.mode == "word":                pdf2word(pdf, out_file)            else:                pdf2excel(pdf, out_file)            print(f"[{i}/{len(pdfs)}] ✓ {pdf.name}")        except Exception as e:            print(f"[{i}/{len(pdfs)}] ✗ {pdf.name} -> {e}")if __name__ == "__main__":    main()
保存为 pdf_convert.py,用法:
# 转 Wordpython pdf_convert.py ./pdfs -o ./docs -m word# 转 Excelpython pdf_convert.py ./pdfs -o ./sheets -m excel

六、几个必须提醒的坑

1. 扫描版 PDF 转不了

pdf2docx 和 pdfplumber 处理的都是"文字型 PDF"。如果你的 PDF 是扫描件(本质是一张张图片),提取出来是空的。

这种情况需要先做 OCR:

pip install ocrmypdfocrmypdf -l chi_sim input.pdf output.pdf

或者用 PaddleOCR、百度 OCR API 处理。识别完再走上面的流程。

2. 复杂排版会错位

多栏排版、图文混排、浮动文本框,转成 Word 后大概率会乱。这是所有转换工具的通病,没有完美方案。建议转换后人工抽查几份。

3. 表格无边框时提取困难

pdfplumber 默认按线条识别表格。如果表格没有边框线,需要改用文字位置来推断:

table = page.extract_table({    "vertical_strategy": "text",    "horizontal_strategy": "text",})
或者换用 camelot 的 stream 模式:
import camelottables = camelot.read_pdf("file.pdf", pages="all", flavor="stream")tables.export("out.xlsx", f="excel")

lattice 模式适合有框线的表格,stream 适合无框线的。

4. 加密 PDF

带密码的文件需要先解密:

pdfplumber.open("file.pdf", password="123456")

pdf2docx 则需要先用 pikepdf 解密再转换。

5. 合并单元格

PDF 里的合并单元格,提取后通常会变成"只有第一行有值,其余为空"。需要在 pandas 里补全:

df["部门"] = df["部门"].replace("", pd.NA).ffill()

七、总结

整个方案的核心就两个库:

  • pdf2docx:PDF → Word,保留版式

  • pdfplumber + pandas:PDF → Excel,结构化数据

适用场景:格式规整的电子版 PDF(合同、报表、发票、名单)。

不适用:扫描件(先 OCR)、超复杂排版(认命手动吧)。

相关学习资料