ARTICLE · 1130657
DeepSeek-Harness 源码深读(17):并发写同一个文件,输的那个凭什么知道自己输了
摘要:写文件表面是 readFile 加 writeFile 的事,fs-local 用了一千两百行:realpath 给文件发身份证,一条 promise 链锁住读改写窗口,五元组版本 token 判过期,发布走 hard link、ReplaceFileW、rename 三条原子路。输掉并发竞争的写者拿到 FS_STALE_VERSION,错误码本身就是补救指令。
模型改一个文件,流程看着很朴素:readText 拿内容,改一行,editText 写回去。两次系统调用的事,dsh 却给了它一个 1210 行的包,packages/fs/fs-local,其中 fsio.ts 781 行、index.ts 265 行、win32.ts 134 行。行数不是重点,重点是磁盘是个并发的世界:你读文件的功夫,另一个 agent 也在写同一个文件;一个文件可以有软链别名;编辑器的 format-on-save 会 touch 你的 mtime;Windows 上一次替换就能把 ACL 弄丢;写到一半还会崩溃。我们跟着一次 editText 走,看每个坑各自长出哪段代码。
先把位置钉住。fs 这条链是四环:dsh-fs 定义抽象,fs-local 落到真实磁盘,fs-observation-policy 管观察策略,tool-fs 把读写编辑包成模型工具。本篇拆第二环。还有个容易误会的地方:配置里的 cwd 只是相对路径的解析基准,不是访问边界,index.ts:59-63 的注释原文写的是 a resolution default, NOT a containment boundary。默认部署里真正挂上 ctx.fs 的其实是它的子类 SandboxedFileSystem(fs-sandbox/src/index.ts:54-56),围栏判断外包给了 ctx.sandboxPolicy,读写发布这些写路径全部继承本篇要拆的基类。
一、你读的,还是刚才那份吗
主线案例开场。模型 readText 读 config.ts,工具层顺手记下这个文件的版本串,形如 4:987654:812:1759237448123456789:1759237449987654321。然后模型改了一行,editText 带着 old_string、new_string 和刚才那个版本串回来。
危险在中间这段空档。readText 到 editText 之间,文件可能被人动过:另一个 agent 覆写了它,或者只是被 touch 了一下。这时编辑继续会发生什么?基于旧内容打补丁,补丁打进新内容,轻则 FS_EDIT_NOT_FOUND,重则语义损坏。反过来静默覆盖也不行,别人的修改直接没了。
要的结果是第三种:输掉竞争的写者拿到一个明确的错误,错误里带着补救指令。而且这个判定要确定,同样的事件序列跑两遍,结果一样。fs-local 的全部设计可以一句话概括:用身份证、锁、版本号三件套,把这件事做确定。我们先看身份证。
这次旅程的全貌一张图:

三件套接住并发
注意看判决那一格:过期检查排在 literal 匹配前面,每个错误码旁边跟着的补救动作都不一样。后面几节把图上的格子逐个拆开。
二、一个文件,两个名字
软链是第一个坑。假设仓库里 ln -s config.ts current.ts,现在同一个文件有两个路径。如果并发控制按路径字符串记账,current.ts 和 config.ts 就是两个文件:一边的版本观察对另一边不可见,穿软链写还可能把链接本身替换掉。
fs-local 的答案是给每个目标两套身份表示,这是本包的核心设计:
export interface LocalTarget { /** Absolute path (symlinks not resolved) — used for display. */ displayPath: string /** Realpath identity — used as the stable target key and the I/O path. */ targetKey: FsTargetKey}这段是身份的类型定义(fsio.ts:106-111)。注意看两个注释:displayPath 给人看,符号链接不解析;targetKey 是 realpath 身份,既当并发控制的稳定键,也当实际 IO 走的路径。身份怎么来的:
const displayPath = resolve(cwd, path)try { // Prefer the file's own realpath (resolves a symlinked file to its target). return { displayPath, targetKey: FsTargetKey(await realpath(displayPath)) }} catch (error: unknown) { // … if (isENOTDIR(error)) throw new FsError(`cannot resolve 「${displayPath}」: a parent path segment is not a directory`, 'FS_NOT_FOUND') if (!isENOENT(error)) throw error}// File absent: realpath the nearest existing ancestor and re-append the// missing suffix (the file basename plus any not-yet-created intermediate// dirs), so the key is stable across creation of those dirs.const missing = [basename(displayPath)]let ancestor = dirname(displayPath)while (true) { try { const realAncestor = await realpath(ancestor) // On Windows, realpath of a regular file succeeds where POSIX returns // ENOTDIR (the OS reports ENOENT for `regular-file/child`, not ENOTDIR). // Stat the ancestor to restore the semantic distinction: a non-directory // ancestor means the target passes through a file and can never be created. if (process.platform === 'win32') { const parentInfo = await stat(realAncestor) if (!parentInfo.isDirectory()) { throw new FsError(`cannot resolve 「${displayPath}」: a parent path segment is not a directory`, 'FS_NOT_FOUND') } } return { displayPath, targetKey: FsTargetKey(join(realAncestor, ...missing)) } } catch (error: unknown) { // … }}这段是解析函数 resolveLocalTarget(fsio.ts:146-194,中间省略两处错误分支)。注意看它处理的是文件还不存在的场景:realpath 最近的现存祖先,再把缺失后缀拼回去。为什么要绕这一圈?创建目录前后的 targetKey 必须一致,否则 stale guard 在创建路径上就失效了,先写后建的目录会让同一次写的前后身份对不上。Windows 那个 if 也在补语义:POSIX 对 regularfile/child 报 ENOTDIR,Windows 报 ENOENT,祖先回溯会把它当缺失路径继续走,所以额外 stat 一次祖先,发现不是目录就抛 FS_NOT_FOUND,手工把 POSIX 的区分补回来。
两套身份的收益,模块头注释写得很直白(index.ts:2-3):别名共享过期守卫,软链和它的目标同一个 targetKey,观察记录互通;穿软链的写更新目标文件,而不是替换链接。
三、一条 promise 链,把读改写关进单间
身份定了,下一个坑是时序。writeText 内部要跑一串动作:probe 现状、比对版本、读 diff 基线、发布。这个窗口如果没有护栏,两个并发写会交错,版本比对就成了笑话,比的是几毫秒前的旧照片。
护栏是一段十几行的 promise 链(index.ts:91-104):
private async withLock<T>(targetKey: string, op: () => Promise<T>): Promise<T> { const prior = this.locks.get(targetKey) ?? Promise.resolve() const run = prior.then(op, op) // Keep the chain alive but swallow this op's result/throw for the *next* waiter. const tail = run.then(() => undefined, () => undefined) this.locks.set(targetKey, tail) try { return await run } finally { if (this.locks.get(targetKey) === tail) { this.locks.delete(targetKey) } }}这段是变更操作的互斥器。注意看第三行的 tail:把 run 的结果和异常都吞掉,只留一个已经落定的 promise 给下一个等待者接续,前一个操作抛没抛异常,后一个照常执行。锁按 targetKey 分道,软链别名在这里也汇进同一间。队列排空就删键(finally 里那两行),Map 不会越养越大。
为什么不用平台文件锁?锁注释自己给了答案(index.ts:74-77):按 targetKey 把变更操作排成 FIFO,让 read→guard→write 窗口不被插入,并发写的结果是确定的,一个成功,其余的看到新版本后以版本过期为由被拒绝。平台文件锁跨平台行为不一致,死锁恢复又是新复杂度;promise 链只在单进程内有效,但 dsh 的写竞争本来就发生在同一进程的插件之间,换来的是完全确定的语义。
四、版本 token:过期是一门手艺
串行解决了插队,没解决另一个问题:排到你的时候,文件还是不是模型读的那份。判据是版本串,生成只要一行(fsio.ts:74-76):
/** Opaque version token from high-resolution identity and freshness metadata. */function versionOf(info: BigIntStats): FsVersion { return FsVersion(`${info.dev}:${info.ino}:${info.size}:${info.mtimeNs}:${info.ctimeNs}`)}这段是版本的配方。注意看五元组各管什么:dev 加 ino 是设备号加 inode,文件的稳定身份,mtime 被重置也认得出它;size 加 mtimeNs 加 ctimeNs 是新鲜度,内容变过没有。BigInt stat 保住纳秒精度,两个快速连续的写也能区分。
写路径的守卫顺序(index.ts:172-187):
const existing = await probe(target.targetKey)if (existing && existing.type !== 'file') { throw new FsError(`cannot write 「${target.displayPath}」: not a regular file`, 'FS_NOT_REGULAR_FILE')}if (expected?.kind === 'replaceIfVersion') { // Stale guard: the file must still exist at the version the owner observed. if (!existing) throw new FsError(`cannot write 「${target.displayPath}」: file no longer exists`, 'FS_STALE_VERSION') if (existing.version !== expected.version) { throw new FsError(`cannot write 「${target.displayPath}」: file changed since it was read`, 'FS_STALE_VERSION`) }} else if (expected?.kind === 'createIfAbsent' && existing) { // createIfAbsent onto an existing file: a blind overwrite — require a read first. throw new FsError(`cannot overwrite existing 「${target.displayPath}」 without reading it first`, 'FS_NOT_OBSERVED')}这段是 writeText 进锁后的前三道判决。注意看错误码的分工:文件没了、版本不一致,都是 FS_STALE_VERSION;想创建但文件已存在,是 FS_NOT_OBSERVED,盲覆盖必须先读。editText 的顺序更讲究(index.ts:227-239):
const existing = await probe(target.targetKey)// Stale guard before literal matching: an edit based on an old read reports// FS_STALE_VERSION, not FS_EDIT_NOT_FOUND/FS_AMBIGUOUS_EDIT against newer content.// Missing targets use the same stale code on guarded and unconditional edit paths.if (!existing) throw new FsError(`cannot edit 「${target.displayPath}」: file changed since it was read`, 'FS_STALE_VERSION')// …if (expected && existing.version !== expected.version) { throw new FsError(`cannot edit 「${target.displayPath}」: file changed since it was read`, 'FS_STALE_VERSION`)}这段是编辑路径的过期判决。注意看注释的第一句:版本检查排在 literal 匹配之前。顺序换过来会怎样?基于旧读的编辑会撞上新内容,报的错是 FS_EDIT_NOT_FOUND,模型得到的补救提示是去改 old_string,可它真正该做的是重新读文件。错误码在这里就是路由指令:FS_STALE_VERSION 告诉模型你的观察过期了,重读;FS_EDIT_NOT_FOUND 告诉模型字面量没匹配上,改匹配串。两种失败,两种补救,模型不用猜。
版本串也解释了主线案例里 format-on-save 的下场:哪怕内容一个字没动,touch 改了 mtimeNs,五元组就变了,旧的观察记录判过期。偏保守,换来的是不漏判。
五、发布:三条原子路
守卫都过了,最后一步是把新内容放到目标路径上。直接 writeFile 会留半截文件,崩溃恢复时无从判断写到哪了。fs-local 的发布走原子三选一(fsio.ts:577-595):
if (createIfAbsent !== undefined) { try { await linkFile(tempPath, absolutePath) } catch (error: unknown) { await throwGuardedCreateFailure(error, absolutePath, createIfAbsent.displayPath, inspectPublicationTarget) }} else if (platform === 'win32' && mode !== undefined) { try { await replaceFile(absolutePath, tempPath) } catch (error: unknown) { // If the observed target disappears during staging, the protected DACL // already copied to the temp remains authoritative for recreation. if (!isENOENT(error)) throw error await rename(tempPath, absolutePath) }} else { await rename(tempPath, absolutePath)}这段是发布的分岔口。注意看三条路各自的语义:创建场景用 hard link,链接已存在会 EEXIST,这是操作系统级的 no-replace 原语,两个并发创建者只有一个能赢;Windows 替换场景用 ReplaceFileW,保住 ACL 等替换元数据;其余场景 rename,同卷原子。
link 失败后的处理值得多看一眼。throwGuardedCreateFailure(fsio.ts:481-516)不直接抛,而是事后探测目标:并发创建者已经放了文件,报 FS_NOT_OBSERVED;目标不是 regular file,报 FS_NOT_REGULAR_FILE;才轮到真 IO 错误。为什么要再探测一次?EEXIST 的 errno 在不同平台、不同文件系统上含义有漂移,按 errno 分类不可靠,看磁盘现状才可靠。
发布前还有一段铺垫(fsio.ts:542-574):mkdir -p 父目录;建私有 staging 目录 .{name}.{pid}.{uuid}.tmpdir,0o700 加 chmod;open(temp, 'wx', 0o600) 独占创建,文件名撞车直接失败;写内容、fsync、还原 mode、close。发布后清理 staging 失败被吞掉(fsio.ts:596-600),注释只给了一句理由:目标文件已经提交成功,残留一个只有 owner 可读的临时目录,不应该让这次写被判定为失败。
六、Windows 的 ACL:替换时悄悄没掉的东西
第六节是纯 Windows 视角。事实是:新建的 Windows 文件会继承所在目录的 DACL。staging 目录是刚建的,temp 文件在里面创建,继承的就是 staging 目录的描述符。直接把 temp rename 到目标位置,目标文件的权限就换了主人,某些部署里这等于权限提升。
处理分两步。第一步在 temp 还是空文件时就复制目标的 DACL(fsio.ts:567-569):
if (platform === 'win32' && mode !== undefined) { await copyFileDacl(absolutePath, tempPath)}这段是替换场景的安全预处理。复制的实现(win32.ts:108-115):
export async function copyFileDaclWin32(source: string, destination: string): Promise<void> { const descriptor = await readFileDaclWin32(source) const api = await win32() const information = (DACL_SECURITY_INFORMATION | PROTECTED_DACL_SECURITY_INFORMATION) >>> 0 if (api.setFileSecurityW(toNamespacedPath(destination), information, descriptor) === 0) { throw win32Error('SetFileSecurityW', api.getLastError(), destination) }}这段是把现有文件的 DACL 抄到 temp 上。注意看 information 那行拼了两个标志:DACL_SECURITY_INFORMATION 是抄描述符,PROTECTED_DACL_SECURITY_INFORMATION 是关键,置上它,目标不再从父目录(staging 目录)继承 ACE。少了这个位,抄过去的描述符会和继承来的混在一起。
第二步靠 ReplaceFileW 在发布时保留替换元数据(win32.ts:122-134)。第五节代码里那个 catch 也在这条线上:目标文件在 staging 期间被人删了,ReplaceFileW 报 ENOENT,退回 rename,注释说已经复制到 temp 的受保护 DACL 仍是权威的,重建出来的文件权限不丢。
Win32 绑定本身是 koffi 懒加载(win32.ts:46-58),advapi32 和 kernel32 只在第一次用到时打开,非 Windows 进程永远不会碰这两个库。
七、两个边角:diff 基线与读上限
主线走完,还有两处细节,是我核稿时多看两眼才咂摸出来的。
一个是行尾。writeText 返回的 after 是 LF 归一化的内容(index.ts:213-217 注释),before 也是,两边共享 diff 基线。不用归一化会怎样?用 CRLF 内容覆盖一个 LF 文件,diff 显示每一行都变了,满屏红。行尾三件套在 fsio.ts:628-649:normalize 把 CRLF 折成 LF,孤立 \r 不动;detect 拿前 4096 字节数 CRLF 和 LF 的个数定风格;restore 在写回时还原,CRLF 分支先重归一化再 join,防止已经是 CRLF 的序列翻倍成 \r\r\n。模型看到的 diff 干净,存储层的行尾风格原样保留。
另一个是 diff 基线的读法(fsio.ts:695-746)。readTextForDiff 是 best-effort 预读,核心两处:
const info = await handle.stat()// …if (info.size >= maxBytes) return nullopenedSize = info.size// One extra byte detects growth after stat without retaining per-read backing buffers.buffer = Buffer.allocUnsafe(openedSize + 1)这段是在已打开的描述符上做尺寸检查。注意看两个点:stat 的是 handle 而不是路径,预检之后文件被人整个换掉,读的还是原来那个描述符;分配 size+1 字节,多读一个字节探测增长,stat 之后文件又变长,读到超出 openedSize 就返回 null,缓冲区不会无上限。结尾的 catch 也讲道理:描述符期的任何 errno 都返回 null,已经提交的写不能因为一次只为生成 diff 的预读失败而被判定为失败;但取消是调用方的意图,FsError 照样往上传播。
八、账单
三件套各自管了什么?双身份把软链别名汇进同一本账,promise 链关掉插队窗口,版本 token 把过期变成可判定。合起来,并发写同一个文件的结果跑两遍都一样:一个成功,其余的拿到 FS_STALE_VERSION。全程没有平台文件锁,也没有一行事件代码,fs/write-intent、fs/observed 这些观察事件全部由 tool-fs 派发,本包只消费 FsWriteIntent 这类数据契约,服务和策略分得很干净。
代价呢?realpath 每次解析都是系统调用;staging 目录每写一次建一删一次,小文件高频写有可测开销;mtimeNs 进版本串,touch 一下旧观察就作废,对 format-on-save 这类场景偏保守;NFD/NFC 这类 Unicode 归一形式不算进身份模型,两个视觉相同的路径仍是两个键。每一条都有注释认账,没有一条藏着。
回到主线案例。输掉的那个写者拿到的错误,就是 index.ts:182 那行 file changed since it was read,手里攥着过期的 4:987654:812:...,下一次 readText 从新版本起步。赢家那边,diff 基线是 LF 归一化的,落在还开着的 turn 里。
输家知道自己怎么输的,靠版本串;那观察记录怎么登记、哪些写要先过用户批准、策略在哪一层生效?这些在 fs-observation-policy,下一篇拆它。
你的 Agent 写文件时,并发写同一文件是怎么处理的?评论区聊聊。
本系列基于 DeepSeek Harness 源码(MIT,0.1.1-rc.1)与官方 Agent Notes 整理,仓库:github.com/deepseek-ai/deepseek-harness。有收获就点个关注,下一篇见。