夜雨聆风学习资料网

ARTICLE · 1114808

Reranker 安装与验证教程

Reranker 安装与验证教程

1. 作用与技术链路

本文记录 Windows + WSL2 + CPU 环境下安装和验证中文 Reranker 的过程。

本教程使用模型:

BAAI/bge-reranker-v2-m3

Reranker 用于对初次检索得到的候选知识片段进行二次排序,不负责生成答案,也不能替代 Embedding:

问题 -> Embedding 召回候选片段 -> Reranker 二次排序 -> 选择最相关片段 -> Ollama 生成回答

2. 环境与资源

项目配置
系统Windows + WSL2
Python3.12.3
推理设备CPU
虚拟环境~/venvs/knowledge-reranker
模型BAAI/bge-reranker-v2-m3
最大输入长度512

模型文件约 2.3GB,缓存实际约 2.2GB。模型加载到 CPU 时会额外占用数 GB 内存;Python 进程退出后,运行内存会释放,缓存文件仍保留在磁盘。

3. 创建 WSL Python 虚拟环境

以下命令必须在 WSL 终端执行,不要在 Windows PowerShell 中执行。

sudo apt updatesudo apt install -y python3.12-venv python3-pippython3 -m venv --clear ~/venvs/knowledge-rerankersource ~/venvs/knowledge-reranker/bin/activate

检查环境:

which pythonpython --version

正常结果类似:

/home/longsanson/venvs/knowledge-reranker/bin/pythonPython 3.12.3

重新打开 WSL 后,需要再次激活:

source ~/venvs/knowledge-reranker/bin/activate

4. 安装 CPU 依赖

保持虚拟环境激活,在 WSL 执行:

python -m pip install --upgrade pippython -m pip install torch \--index-url https://download.pytorch.org/whl/cpupython -m pip install sentence-transformers

检查依赖:

python - <<'PY'import torchimport sentence_transformersprint("Torch版本:", torch.__version__)print("Sentence Transformers版本:", sentence_transformers.__version__)print("CUDA可用:", torch.cuda.is_available())PY

没有 NVIDIA GPU 时,CUDA可用:False 是正常结果。

5. Hugging Face 网络与代理

第一次加载模型会从 Hugging Face 下载文件。先在 WSL 测试:

curl -4 -I --connect-timeout 10 https://www.baidu.comcurl -4 -I --connect-timeout 10 https://huggingface.co

如果普通网站正常而 Hugging Face 超时,可以使用 Clash Verge。

5.1 Clash 设置

在 Clash Verge 中开启:

系统代理:开启局域网连接:开启端口:7897

在 Windows PowerShell 中验证本机代理:

curl.exe -I -x http://127.0.0.1:7897 --connect-timeout 15 https://huggingface.co

看到 HTTP/1.1 200 Connection established 和 HTTP/2 200,说明 Clash 正常。

5.2 WSL 使用 Clash

在 WSL 获取 Windows 宿主机地址:

WIN_HOST=$(ip route | awk '/default/ {print $3}')echo $WIN_HOST

设置当前终端的临时代理:

export HTTP_PROXY="http://${WIN_HOST}:7897"export HTTPS_PROXY="http://${WIN_HOST}:7897"export ALL_PROXY="http://${WIN_HOST}:7897"

通过代理测试:

curl -I --proxy "http://${WIN_HOST}:7897" \  --connect-timeout 15 \  https://huggingface.co

返回 HTTP/2 200 或 HTTP/1.1 302 后,再运行模型脚本。

5.3 WSL 访问代理端口超时

如果 Windows 本机代理正常、WSL 访问 7897 超时,可以在管理员 PowerShell 中允许当前 WSL 子网访问:

New-NetFirewallRule `  -DisplayName "Clash-Wsl-7897" `  -Direction Inbound `  -Action Allow `  -Protocol TCP `  -LocalPort 7897 `  -RemoteAddress 192.168.128.0/20 `  -Profile Any

模型下载完成后,清除当前 WSL 终端的代理:

unset HTTP_PROXYunset HTTPS_PROXYunset ALL_PROXY

不再需要防火墙规则时,在管理员 PowerShell 中删除:

Remove-NetFirewallRule -DisplayName "Clash-Wsl-7897"

6. 加载和验证模型

保持虚拟环境激活,执行:

python - <<'PY'import timeimport torchfrom sentence_transformers import CrossEncoderprint("Torch版本:", torch.__version__)print("CUDA可用:", torch.cuda.is_available())load_start = time.time()model = CrossEncoder(    "BAAI/bge-reranker-v2-m3",    max_length=512,    device="cpu")load_seconds = time.time() - load_startpairs = [    (        "普通员工出差乘坐高铁应该选择什么席别?",        "普通员工出差乘坐高铁原则上选择二等座。"    ),    (        "普通员工出差乘坐高铁应该选择什么席别?",        "项目验收完成后,项目负责人应提交验收报告。"    ),    (        "差旅报销应该什么时候提交?",        "差旅报销应在行程结束后十五个工作日内提交。"    )]predict_start = time.time()scores = model.predict(pairs)predict_seconds = time.time() - predict_startprint("Reranker模型加载成功")print("模型加载耗时:", round(load_seconds, 2), "秒")print("排序耗时:", round(predict_seconds, 2), "秒")for index, score in enumerate(scores, start=1):    print(f"第{index}组分数:{float(score):.6f}")PY

本次 CPU 验证结果:

Torch版本: 2.14.0+cpuCUDA可用: FalseReranker模型加载成功模型加载耗时: 191.21 秒排序耗时: 0.32 秒第1组分数:0.998868第2组分数:0.000016第3组分数:0.995958

结果说明:第 1、3 组是相关问题和资料,分数高;第 2 组无关,分数低。Reranker 分数用于同一批候选片段的相对排序,不应直接当作概率或百分比。

7. 模型缓存与运行内存

模型默认缓存位置:

/home/longsanson/.cache/huggingface/hub

查看缓存大小:

du -sh ~/.cache/huggingface/hub

查看模型目录:

find ~/.cache/huggingface/hub \  -maxdepth 1 \  -type d \  -iname '*bge*'

模型文件会持续占用磁盘;只有加载模型的 Python 进程运行时才会占用主要内存。退出脚本或执行下面命令会释放当前 Shell 环境,但不会删除模型:

deactivate

8. 常见问题

8.1 虚拟环境创建失败

如果出现 The virtual environment was not created successfully,执行:

sudo apt updatesudo apt install -y python3.12-venv python3-pippython3 -m venv --clear ~/venvs/knowledge-reranker

8.2 python: command not found

说明虚拟环境没有激活:

source ~/venvs/knowledge-reranker/bin/activate

8.3 Network is unreachable 或 Hugging Face 超时

这是模型下载阶段的网络问题,不是 Reranker 代码错误。先测试:

curl -I --connect-timeout 10 https://huggingface.co

必要时按本文第 5 节配置 Clash。不要同时启动多个下载脚本。

8.4 PowerShell 中执行 Linux 路径报错

以下命令必须在 WSL 执行:

du -sh ~/.cache/huggingface/hubfind ~/.cache/huggingface/hub -maxdepth 1 -type d

如果必须在 PowerShell 中调用 WSL:

wsl -e bash -lc "du -sh ~/.cache/huggingface/hub"

9. 接入智能知识库

推荐的第一阶段参数:

候选召回:20~50 个片段Reranker 最终选择:5~8 个片段max_length:512CPU 并发:1

完整流程:

权限过滤 -> 关键词检索 + 向量检索 -> 召回 20~50 个候选片段 -> Reranker 排序 -> 选择前 5~8 个互补片段 -> Ollama 生成带引用答案

权限过滤必须发生在召回前,不能先把无权限资料交给 Reranker 或大模型。

正式接入时,建议将 Reranker 封装为独立的 Python 服务 knowledge-reranker,业务服务通过内部 HTTP 接口调用。这样后续可以在不修改业务代码的情况下切换 GPU、vLLM 或其他排序模型。

相关学习资料