夜雨聆风学习资料网

ARTICLE · 1118169

Codex 源码-System Prompt 组装与注入

Codex 源码-System Prompt 组装与注入

Codex 源码解析系列

第 9 讲:System Prompt 组装与注入

基于 OpenAI Codex 源码 · 2026-10-03

💡 本讲一句话:Codex 的"系统提示词"不是一句写死的字符串,而是一条流水线:会话启动时按三级优先级解析 base instructions(配置覆盖 → 历史继承 → 模型模板),之后每一轮再把沙箱策略、审批规则、AGENTS.md、插件能力等十几个"世界状态分节"做快照对比,只把变化的部分渲染成 developer/user 消息注入上下文。读完你会明白一个 agent 的提示词是怎么被"动态拼装 + 差量维护"的。

一、先建立地图:提示词从哪来、往哪去

上一讲我们把 Session / TurnContext / StepContext 三层快照拆完了。这一讲回答一个更基础的问题:模型每次看到的"系统提示词"到底是谁、在什么时候、按什么规则拼出来的?

源码里这条管线横跨四个 crate,但骨架只有四步:

🔹 解析:会话启动时确定 base instructions(core/src/session/mod.rs)🔹 分节:把沙箱、审批、AGENTS.md、插件等状态组织成 WorldState sections(core/src/context/world_state/)🔹 差量渲染:每轮对比快照,只输出变化的 fragment(core/src/context_manager/updates.rs)🔹 组装请求:fragment 合并成消息,按 provider 能力走两条不同的 API 路径(core/src/client.rs)

下面按这条管线逐段拆源码。

二、Base instructions 的三级优先级链

"base instructions"是 Codex 对系统提示词的正式称呼。它在哪解析?core/src/session/mod.rs 的会话初始化路径里,注释把优先级写得明明白白:

📄 codex-rs/core/src/session/mod.rs (第 687-722 行,节选)

// Resolve base instructions for the session. Priority order: // 注释直接声明三级优先级// 1. config.base_instructions override // 第一级:用户显式配置,最高权威// 2. conversation history => session_meta.base_instructions // 第二级:从历史会话继承(resume/fork)// 3. rendered instructions_template for current model // 第三级兜底:渲染当前模型的模板let base_instructions = config // 从配置层开始找    .base_instructions // 用户显式覆盖项(config.toml / API 参数传入)    .clone() // clone 出来,避免长期借用 Arc<Config>    .or_else(|| conversation_history.get_base_instructions().map(|s| s.text)) // 没配置 → 查历史 session_meta:resume/fork 时继承父线程的指令    .unwrap_or_else(|| model_info.get_model_instructions(config.personality)); // 都没有 → 渲染当前模型模板(personality 参与渲染)

为什么这样设计:三级链的顺序就是"意图强度"排序——用户亲手写的配置 > 会话历史里已经生效的指令(resume/fork 必须保持行为连续,否则换台机器继续对话 agent 就"失忆变人")> 模型自带的默认模板。注意第三级不是读一个静态文件,而是 get_model_instructions(personality)——指令是按模型渲染的,不同模型可以带不同的系统提示词。这个入口在 protocol crate:

📄 codex-rs/protocol/src/openai_models.rs (第 534-552 行)

pub fn get_model_instructions(&self, personality: Option<Personality>) -> String { // 渲染模型指令的入口;personality 可选    if let Some(model_messages) = &self.model_messages // 先看该模型有没有服务端下发的 messages 配置        && let Some(template) = &model_messages.instructions_template // 再取其中的指令模板(let-chains 短路)    {        if model_messages.instructions_variables.is_none() { // 没声明变量 → 模板就是纯文本,原样返回            return template.clone();        }        let personality_message = model_messages // 有变量 → 按当前 personality 查对应文案            .get_personality_message(personality)            .unwrap_or_default(); // 查不到就用空串(宁可少说,不报错)        template.replace(PERSONALITY_PLACEHOLDER, personality_message.as_str()) // 把 {{ personality }} 占位符替换成人格化文本    } else {        warn!(model = %self.slug, "Model has no instruction template; returning empty instructions."); // fail loud:打警告日志而不是静默塞错内容            String::new() // 返回空指令,让上层决定怎么办    }}

为什么这样设计:模板 + 占位符({{ personality }})是"一份模板、多种人格"的标准做法——模型方可以在服务端改文案而不用发版客户端。两个细节值得注意:一是 instructions_variables.is_none() 时直接当字面文本返回,说明"有模板但没变量"是合法状态;二是缺模板走 warn! + 空串——fail loud,宁可让会话没有系统提示词并留下日志,也不悄悄用别的模型的文案顶替。

默认模板长什么样?仓库里有一份 protocol/src/prompts/base_instructions/default.md(275 行),开头就是:

📄 codex-rs/protocol/src/prompts/base_instructions/default.md (第 1-6 行)

You are a coding agent running in the Codex CLI, a terminal-based coding assistant. // 第一句定身份:Codex CLI 里的编码智能体Your capabilities: // 能力清单开头- Receive user prompts and other context provided by the harness, such as files in the workspace. // 输入侧:用户提示 + harness 提供的上下文(如工作区文件)- Communicate with the user by streaming thinking & responses, and by making & updating plans. // 输出侧:流式思考/回复 + 制定和更新计划- Emit function calls to run terminal commands and apply patches. // 行动侧:发函数调用跑命令、打补丁

注意模板里还内嵌了 AGENTS.md 规范("更深层嵌套的 AGENTS.md 优先")、preamble 消息风格、沙箱与审批说明——也就是说系统提示词本身就在教模型怎么和 harness 协作,而不只是描述任务。这里有个容易忽略的细节:模板里提到 update_plan(计划工具)的段落,在功能关闭时会被手术式切除:

📄 codex-rs/core/src/session/mod.rs (第 1399-1417 行)

/// Render the request copy without changing instructions persisted or inherited by forks. // 关键设计:只改"本次请求的副本",不动持久化状态pub(crate) async fn get_prompt_base_instructions(&self) -> BaseInstructions {    let config = self.get_config().await; // 取当前配置    let instructions = self.get_base_instructions().await; // 读会话级 base instructions(含 provenance 来源标记)    if !config.update_plan_enabled && config.model_catalog.is_none() && matches!(instructions.provenance, Some(BaseInstructionsProvenance::Model { .. })) { // 三个条件同时满足才动刀:计划工具关闭 + 无自定义目录 + 指令确实来自模型模板        BaseInstructions { text: crate::context::without_update_plan_instructions(&instructions.text), ..instructions } // 切除 update_plan 段落,其余字段原样保留    } else { instructions } // 用户自定义的指令绝不修改——尊重用户意图}

为什么这样设计:这是全篇最精巧的一处。切除逻辑(core/src/context/update_plan_instructions.rs)按行扫描,只删 ## Planning / ## `update_plan` 等标题到下一个同级标题之间的段落——其余字节原样保留。而触发条件里 provenance == Model 这一条是护栏:只有"指令确实来自模型模板"才允许动,用户自己写的 base instructions 一个字都不碰。同时它只作用于请求副本(函数名里的 "prompt"),持久化到 rollout 的原文不变——fork/resume 时继承的仍是完整版本,行为可复现。

优先级
来源
行为特征
1(最高)
config.base_instructions
用户显式覆盖;provenance=Custom,任何自动改写逻辑都跳过它
2
conversation history → session_meta
resume/fork 继承父线程指令,保证跨会话行为连续
3(兜底)
model_info.get_model_instructions()
按模型渲染模板 + personality 占位符;缺模板时 warn + 空串

三、ContextualUserFragment:所有注入的统一协议

base instructions 只是"静态底座"。真正让提示词活起来的是动态注入:沙箱策略变了要告诉模型、用户 @ 了插件要提醒模型、换模型后要重新交代规则……这些内容五花八门,Codex 用一个 trait 把它们统一成同一种东西——ContextualUserFragment(context-fragments crate):

📄 codex-rs/context-fragments/src/fragment.rs (第 64-119 行,节选)

pub trait ContextualUserFragment { // 统一注入协议:任何"上下文片段"实现它就能进入模型输入    fn role(&self) -> &'static str; // 以哪个角色说话(developer/user)——决定消息归属    /// Returns a stable `<feature>.<name>` classification, using `generic` for shared fragments. // content_kind:稳定分类标识,用于去重/审计/遥测    fn content_kind(&self) -> ContentItemKind;    /// Whether this fragment must be recorded as its own response item. // 是否必须独占一条消息(默认否 → 可合并)    fn requires_separate_message(&self) -> bool { false }    fn markers(&self) -> (&'static str, &'static str); // 起止标记:之后能在历史里"认出"这个片段(去重/压缩时识别)    fn body(&self) -> String; // 模型可见的正文内容    fn render(&self) -> String { // 渲染 = 标记 + 正文,不额外加分隔符(空白由实现方自己控制)        let (start_marker, end_marker) = self.markers();        let body = self.body();        if start_marker.is_empty() && end_marker.is_empty() { return body; } // 无标记 → 纯文本直出        format!("{start_marker}{body}{end_marker}") // 有标记 → 包一层,如 <permissions instructions>...</permissions instructions>    }    fn into(self) -> ResponseItem where Self: Sized { // 一行转成 API 消息项:fragment 与请求体之间的桥        ResponseItem::from(self.render_fragment())    }}

为什么这样设计:这个 trait 把"注入什么内容"和"怎么进请求"彻底解耦。三个字段各管一件事:role() 决定消息归属(developer 还是 user);content_kind() 是稳定身份标识,写进 internal_chat_message_metadata_passthrough 随消息持久化——压缩历史、审计"模型当时看到了什么"都靠它;markers() 是文本级指纹,让系统能在纯文本历史里重新识别出某个片段(下一节的差量注入全靠它)。看一个最小实现——base instructions 自己也是 fragment:

📄 codex-rs/core/src/context/base_instructions.rs (第 4-31 行)

#[derive(Clone, Debug, PartialEq, Eq)] // 派生比较能力:差量判断需要"两个片段是否相等"pub(crate) struct BaseInstructionsFragment(pub(crate) String); // base instructions 片段:包一层最终指令文本impl ContextualUserFragment for BaseInstructionsFragment { // 实现统一注入协议    fn content_kind(&self) -> ContentItemKind { // 稳定标识 model.base_instructions——历史里去重/审计的钥匙        ContentItemKind("model.base_instructions".to_string())    }    fn role(&self) -> &'static str { "developer" } // developer 角色:系统级指令,不是用户说的话    fn requires_separate_message(&self) -> bool { true } // 必须独占一条消息——绝不和其他上下文混排    fn markers(&self) -> (&'static str, &'static str) { Self::type_markers() } // 标记取类型级默认值    fn type_markers() -> (&'static str, &'static str) { ("", "") } // 空标记 → render 出纯文本,matches_text 永不命中(它靠 content_kind 识别)    fn body(&self) -> String { self.0.clone() } // 正文就是指令文本本身}

为什么这样设计:requires_separate_message = true 是关键——base instructions 必须独占一条 ResponseItem,不能和"当前时间提醒""插件提示"挤在同一条消息里。原因有二:一是可审计性(压缩历史时能整条保留或整条丢弃);二是 client.rs 的 Responses Lite 路径要给它单独派生稳定 ID(第八节细讲)。而它用空标记 + content_kind 识别,说明 Codex 对"身份"有两套机制:结构化元数据优先,文本标记兜底——老版本历史里没有 metadata 时,就靠 <permissions instructions> 这类标记认人。

四、WorldState:把"世界状态"切成可差量的节

如果每轮都把全部上下文重发一遍,token 成本会爆炸。Codex 的解法是把所有"模型可见的世界状态"组织成 WorldState sections——每个 section 自己管快照、自己决定"变了才说什么":

📄 codex-rs/core/src/context/world_state/mod.rs (第 228-262、399-413 行,节选)

pub(crate) trait WorldStateSection: Send + Sync + 'static { // 一个"世界状态节":自管快照、自管差量渲染    const ID: &'static str; // 稳定 ID(持久化进 rollout,跨版本不能变)——section 的身份    type Snapshot: DeserializeOwned + Serialize; // 快照类型:只存"判断变化所需的最小数据"    fn snapshot(&self) -> Self::Snapshot; // 当前状态 → 可序列化快照(每轮持久化,供下轮对比)    /// Whether the section contributes comparison state to persisted rollouts. // 是否参与持久化(默认是;纯展示型节可以关掉)    fn should_persist(&self) -> bool { true }    /// Recognizes legacy fragments whose identity depends on this section's current value. // 识别旧版本片段文本——兼容升级前的历史    fn matches_current_legacy_fragment(&self, role: &str, text: &str) -> bool { Self::matches_legacy_fragment(role, text) }    /// Whether retained model history must still contain this section's rendered fragment. // 是否要检查"片段还在历史里吗"(防压缩后重复注入)    fn has_retained_fragment_matcher() -> bool { false }    fn render_diff( // 核心方法:对比上一份快照,决定本轮告诉模型什么        &self,        previous: PreviousSectionState<'_, Self::Snapshot>, // 三态:Absent(没有)/ Unknown(有但读不出类型)/ Known(精确快照)    ) -> Option<Box<dyn ContextualUserFragment>>; // None = 没变化,零注入;Some(fragment) = 变了,输出差量文本}impl WorldState {    /// Renders every section as new, without any known previous state. // 首轮:所有节都当"新"处理 → 全量注入    pub(crate) fn render_full(&self) -> Vec<Box<dyn ContextualUserFragment>> {        self.render_with(|_, _| PreviousSectionState::Absent) // 统一走 render_with,previous 恒为 Absent    }    /// Renders each section against the exact persisted snapshot when available. // 后续轮:逐节对比持久化快照 → 只输出差量    pub(crate) fn render_diff(&self, previous: &WorldStateSnapshot) -> Vec<Box<dyn ContextualUserFragment>> {        self.render_with(|id, _| match previous.sections.get(id) { // 按 section ID 查上一份快照            Some(previous) => PreviousSectionState::Known(previous), // 有快照 → 精确差量            None => PreviousSectionState::Absent, // 没有(新增的节)→ 当新内容全量注入        })    }}

为什么这样设计:这是典型的"状态机 + 快照对比"模式,和前端框架的 diff 思路同构。三个设计点:① Snapshot 是关联类型——每个 section 自己定义"什么算变化所需的最小状态"(permissions 存哈希+前缀集合,model 只存模型名),而不是统一存一大坨 JSON;② PreviousSectionState 的三态设计承认现实:历史里可能有这个节但快照读不出来(Unknown),此时各 section 自己决定保守策略;③ render_diff 返回 Option——"没变化"是一等公民,零注入不产生任何消息。

看一个最精细的 section:permissions。它的差量策略分三层——指令体没变 + 前缀集合没变 → 什么都不发;指令体没变但新增了已批准命令前缀 → 只发一条"新增这些前缀可用"的短通知;其余情况才重发完整权限说明:

📄 codex-rs/core/src/context/world_state/permissions.rs (第 88-124 行)

fn render_diff( // permissions 节:所有 section 里差量策略最精细的一个    &self,    previous: PreviousSectionState<'_, Self::Snapshot>,) -> Option<Box<dyn ContextualUserFragment>> {    match (previous, &self.snapshot) { // 同时对比"上一状态"和"当前状态"两个快照        (PreviousSectionState::Known(PermissionsSnapshot::Current { instructions: previous_instructions, approved_command_prefixes: previous_prefixes }), PermissionsSnapshot::Current { instructions, approved_command_prefixes }) if previous_instructions == instructions => { // 指令体哈希没变 → 只剩前缀集合可能变了            if previous_prefixes == approved_command_prefixes { return None; } // 前缀也没变 → 本轮零注入(最省路径)            if previous_prefixes.is_subset(approved_command_prefixes) { // 新集合 ⊇ 旧集合 = 纯新增,没有撤销                let added_prefixes = approved_command_prefixes.difference(previous_prefixes).cloned().collect(); // 求差集:本轮新批准的命令前缀                if let Some(prefixes) = format_allow_prefixes(added_prefixes) {                    return Some(Box::new(ApprovedCommandPrefixSaved::new(prefixes))); // 只发一条"这些前缀现在可用了"的短通知,不重发整份策略                }            }        }        (PreviousSectionState::Known(PermissionsSnapshot::Legacy(previous)), PermissionsSnapshot::Current { .. }) if previous == &WorldStateHash::from_fragment(&self.instructions) => return None, // 旧版快照(只有哈希)与当前一致 → 视为没变,兼容升级        _ => {} // 其余情况(策略真变了 / 状态未知)落到下面全量重发    }    Some(Box::new(self.instructions.clone())) // 兜底:重新注入完整 permissions instructions}

为什么这样设计:注意快照里存的是 WorldStateHash(SHA-1)而不是全文——对比用哈希,注入才取原文,持久化体积和对比成本都降下来了。而"纯新增前缀只发短通知"这条路径是真正的省钱点:用户每批准一条命令,上下文里就多一行 ApprovedCommandPrefixSaved,而不是把几百字的权限说明重发一遍。撤销前缀(非子集)则退回全量——因为"哪些被撤了"没法用增量表达得干净。

另一个 section 展示了差量的另一面:换模型。模型变了,旧的系统提示词就失效了,必须重新交代——但只发一次:

📄 codex-rs/core/src/context/world_state/model.rs (第 44-60 行) + model_switch_instructions.rs (第 17-44 行)

fn render_diff( // model 节:快照就是模型名本身(String)    &self,    previous: PreviousSectionState<'_, Self::Snapshot>,) -> Option<Box<dyn ContextualUserFragment>> {    let model_changed = match previous { // 判断"模型是否换了"        PreviousSectionState::Known(previous) => previous != &self.model, // 有精确快照 → 直接比名字        PreviousSectionState::Unknown | PreviousSectionState::Absent => self.previous_model.as_deref().is_some_and(|previous| previous != self.model), // 没快照 → 用运行时记录的"上一轮模型名"兜底判断    };    (model_changed && !self.instructions.is_empty()).then(|| { // 换了且有指令文本才注入(空指令不废话)        Box::new(ModelSwitchInstructions::new(self.instructions.clone())) as Box<dyn ContextualUserFragment> // 包成"换模型通知"片段    })}impl ContextualUserFragment for ModelSwitchInstructions { // 换模型通知:告诉模型"你被换了,按新规则来"    fn content_kind(&self) -> ContentItemKind { ContentItemKind("model_switch.instructions".to_string()) } // 稳定标识    fn role(&self) -> &'static str { "developer" } // developer 角色    fn requires_separate_message(&self) -> bool { true } // 独占一条消息——换模型是大事件,不混排    fn type_markers() -> (&'static str, &'static str) { ("<model_switch>", "</model_switch>") } // XML 式标记:历史里能认出它(防重复注入)    fn body(&self) -> String {        format!("The user was previously using a different model. Please continue the conversation according to the following instructions:\n\n{}\n", self.model_instructions) // 先解释"之前是别的模型",再附完整新指令——给模型一个行为切换的台阶    }}

为什么这样设计:换模型时不是悄悄替换系统提示词,而是显式告诉模型"你之前是另一个模型,现在按这套规则继续"——因为对话历史里还留着旧模型的行为痕迹(它可能引用过旧指令里的约定),直接换文案会让模型困惑。标记 <model_switch> + requires_separate_message 双保险,保证这条通知在历史压缩后仍可被识别、不会被合并进别的消息。

Section ID
快照内容
变化时注入什么
model
模型名(String)
ModelSwitchInstructions:换模型通知 + 完整新指令,独占消息
permissions
指令哈希 + 已批准前缀集合
纯新增前缀 → 短通知;策略变了 → 完整重发;没变 → 零注入
personality
模型 + personality + 是否已烘焙进 base instructions
人格切换时注入对应 personality 文案(若未烘焙)
agents_md / environments / plugins…
各自的最小状态(文件内容哈希、环境选择、插件可用性)
对应 fragment:AGENTS.md 变更、环境切换说明、插件使用说明等
token_budget / context_window_guidance
窗口 ID(first/previous/current)+ guidance 文案
新上下文窗口开启时注入预算提醒与压缩指引

五、Permissions instructions:策略对象 → 提示词文本

permissions section 的"原文"从哪来?prompts/src/permissions_instructions.rs(455 行)负责把策略对象翻译成模型能读懂的英文说明——这是"安全策略即提示词"的关键一环:

📄 codex-rs/prompts/src/permissions_instructions.rs (第 88-118、176-196 行)

impl PermissionsInstructions {    /// Builds permissions instructions from the effective permission profile and approval policy. // 入口:把"策略对象"翻译成"提示词文本"    pub fn from_permission_profile(        permission_profile: &PermissionProfile, // 文件系统 + 网络沙箱策略(真正执行的那份)        approval_policy: AskForApproval, // 审批策略四选一:Never / UnlessTrusted / OnRequest / Granular        approval_context: ApprovalPromptContext<'_>, // reviewer(user/auto_review)+ 模型侧自定义文案覆盖        exec_policy: &Policy, // execpolicy 命令策略引擎:已批准前缀等        cwd: &Path, // 当前工作目录:解析相对可写根用        exec_permission_approvals_enabled: bool, // shell 权限审批流是否开启(影响文案分支)        request_permissions_tool_enabled: bool, // request_permissions 工具是否可用(决定要不要教模型用它)    ) -> Self {        let file_system_sandbox_policy = permission_profile.file_system_sandbox_policy(); // 拆出文件系统策略        let (sandbox_mode, writable_roots) = sandbox_prompt_from_policy(&file_system_sandbox_policy, cwd); // 映射成三档:full access / workspace write / read-only + 可写根列表        Self::from_permissions_with_network_and_denied_reads(            sandbox_mode, // 沙箱档位决定选哪套模板文案            network_access_from_policy(permission_profile.network_sandbox_policy()), // 网络策略 → Enabled/Restricted(填进 {{ network_access }})            PermissionsPromptConfig { approval_policy, approvals_reviewer: approval_context.reviewer, .. }, // 审批相关配置打包            writable_roots, // 可写根列表(渲染成 "The writable root is `...`.")            denied_reads_text(&file_system_sandbox_policy, cwd), // 被禁读的路径/glob → "不要请求升级权限去读它们"的明确警告        )    }}impl ContextualUserFragment for PermissionsInstructions {    fn role(&self) -> &'static str { "developer" } // developer 消息:策略说明属于系统层    fn content_kind(&self) -> ContentItemKind { ContentItemKind("permissions.instructions".to_string()) } // 稳定标识 permissions.instructions    fn type_markers() -> (&'static str, &'static str) { ("<permissions instructions>", "</permissions instructions>") } // 标记:WorldState 靠它认出"这条还在历史里"(has_retained_fragment_matcher=true)}

为什么这样设计:三个值得学的点。① 文案与策略同源:提示词不是手写的,而是从 PermissionProfile(真正执行沙箱的那份对象)推导出来的——模型看到的和系统实际做的永远一致,不存在"提示词说能写、沙箱却拒绝"的漂移。② 模板可被服务端覆盖:approval_messages / permission_messages(来自模型配置)优先于本地 include_str! 模板——OpenAI 可以按模型调文案而不用发版。③ denied reads 单独成段并明确说"不要请求升级权限去读它们,这是策略限制"——直接掐断模型反复试探被禁路径的行为模式。

📌 设计模式小结:到这一节,Codex 提示词系统的三层结构已经完整——静态底座(base instructions,三级优先级解析)+ 动态分节(WorldState sections,快照差量注入)+ 统一协议(ContextualUserFragment,role/kind/markers 三字段)。下一节看两个"事件驱动"的注入点:用户 @ 插件、以及最终的请求组装。

六、Plugin mention injection:@ 了才注入,且带硬预算

插件能力(MCP servers / Apps / skills)默认不进提示词——只有用户这一轮显式提到某个插件,才生成一条 developer hint 指路。入口在 core/src/plugins/injection.rs:

📄 codex-rs/core/src/plugins/injection.rs (第 14-59 行,节选) + render.rs (第 79-88 行)

pub(crate) fn build_plugin_injections( // 插件提及注入入口:用户显式 @ 了插件 → 生成 developer hint    mentioned_plugins: &[PluginCapabilitySummary], // 本轮被显式提到的插件列表(从用户输入里解析)    mcp_tools: &[ToolInfo], // 全部 MCP 工具(用来反查"哪些 server 属于这个插件")    available_connectors: &[connectors::AppInfo], // 可用 App connectors(同样按插件名反查)) -> Vec<ResponseItem> {    if mentioned_plugins.is_empty() { return Vec::new(); } // 没提及 → 零注入:不污染上下文,这是默认路径    mentioned_plugins.iter().filter_map(|plugin| { // 逐个处理被提到的插件        let available_mcp_servers = mcp_tools.iter() // 找出属于该插件的 MCP server(排除内置 apps server)            .filter(|tool| tool.server_name != CODEX_APPS_MCP_SERVER_NAME && tool.plugin_display_names.iter().any(|name| name == &plugin.display_name))            .map(|tool| tool.server_name.clone()).collect::<BTreeSet<String>>() // BTreeSet 去重 + 排序 → 输出顺序稳定(提示词可复现)            .into_iter().collect::<Vec<_>>();        let available_apps = available_connectors.iter() // 找出属于该插件且已启用的 App            .filter(|connector| connector.is_enabled && connector.plugin_display_names.iter().any(|name| name == &plugin.display_name))            .map(connector_display_label).collect::<BTreeSet<String>>()            .into_iter().collect::<Vec<_>>();        render_explicit_plugin_instructions(plugin, &available_mcp_servers, &available_apps) // 渲染 hint 文本(内部带 4KB 硬预算)            .map(PluginInstructions::new).map(ContextualUserFragment::into) // 包成 fragment → ResponseItem,走统一协议    }).collect()}fn bound_explicit_plugin_instructions(rendered: String) -> String { // 硬预算:插件 hint 最多 4KB(MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES)    if rendered.len() <= MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES { return rendered; } // 没超 → 原样通过    let max_prefix_bytes = MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES.saturating_sub(TRUNCATED_PLUGIN_INSTRUCTIONS_SUFFIX.len()); // 先给截断提示语留出空间(saturating 防下溢)    let prefix = take_bytes_at_char_boundary(&rendered, max_prefix_bytes); // 按字符边界截断——绝不把 UTF-8 多字节切半    format!("{prefix}{TRUNCATED_PLUGIN_INSTRUCTIONS_SUFFIX}") // 追加 "Additional plugin capabilities omitted to fit the context limit."}

为什么这样设计:"提及才注入"是按需付费的上下文策略——插件生态可以无限扩张,但提示词预算有限,默认零成本、用到才花钱。两个工程细节:① BTreeSet 而不是 HashSet——插件列表顺序必须稳定,否则同一配置每轮渲染出不同提示词,prompt cache 全废;② 截断用 take_bytes_at_char_boundary——Rust 里按字节切字符串可能产生非法 UTF-8,这个工具函数保证落在字符边界上。截断提示语本身也计入预算(先减后切),这是很多实现会漏的坑。

七、最终组装:fragment 路由 + Prompt 打包

所有 fragment 渲染出来后,build_initial_context_with_world_state(session/mod.rs)负责路由——按角色和标记把它们分进三个桶:

📄 codex-rs/core/src/session/mod.rs (第 4254-4310 行,节选)

for fragment in world_state.render_full() { // 首轮全量渲染 → 按角色 + 标记路由进三个桶    match fragment.role() {        "developer" if fragment.markers().0 == ModelSwitchInstructions::type_markers().0 => { developer_sections.insert(0, fragment.render_fragment()); } // 换模型通知必须插到 developer 上下文最前面(新规则优先于旧约定)        "developer" if fragment.requires_separate_message() && fragment.markers().0.is_empty() => { separate_developer_sections.push(fragment.render_fragment()); } // 要求独占消息的片段 → 各自一条 ResponseItem        "developer" => developer_sections.push(fragment.render_fragment()), // 普通 developer 片段 → 攒进合并桶(最终合成一条大消息)        "user" => contextual_user_sections.push(fragment.render_fragment()), // user 角色片段(如推荐插件)→ 单独的 user 消息        _ => {} // 其他角色忽略    }}let mut items = Vec::with_capacity(4);if let Some(developer_message) = crate::context_manager::updates::build_rendered_message(developer_sections) { items.push(developer_message); } // 合并桶 → 一条 developer 消息(多个 content item)for section in separate_developer_sections { if let Some(m) = build_rendered_message(vec![section]) { items.push(m); } } // 独占片段 → 逐条成消息// ……multi-agent mode / contextual user / guardian policy / managed instructions 依次追加……if separate_guardian_developer_message && let Some(developer_instructions) = turn_context.developer_instructions.as_deref() { /* GuardianPolicy::new(...).render_fragment() → 独立消息 */ } // guardian 子代理的策略提示单独成条——便于审计"审查者当时看到什么"

为什么这样设计:路由规则把"消息粒度"变成了显式决策:能合并的合并(省 API 消息数、保持上下文紧凑),必须独立的独立(base instructions、换模型通知、guardian 策略——各自可审计、可整条压缩)。而 merge_contextual_fragments(context_manager/updates.rs)的合并算法也值得看一眼:它按"相邻 + 同角色 + 都是 Mergeable"三条件把 fragment 拼进同一条消息,遇到 requires_separate_message 就强制断行——顺序敏感、边界清晰。

消息列表就绪后,每轮采样前 build_prompt(turn.rs)把一切打包成 Prompt:

📄 codex-rs/core/src/session/turn.rs (第 1509-1526 行)

pub(crate) fn build_prompt( // 最终打包:历史 + 工具 + base instructions → 一个 Prompt    input: Vec<ResponseItem>, // 对话历史(含所有已注入的上下文片段)    step_context: &StepContext, // 本步冻结快照(工具路由/权限/模型信息都从这里取)    base_instructions: BaseInstructions, // 解析好的 base instructions(带 provenance,调用方刚用 get_prompt_base_instructions 取的请求副本)) -> Prompt {    let turn_context = &step_context.turn;    Prompt {        input, // 历史原样进 input 数组        tools: step_context.tool_router.model_visible_specs(), // 模型可见的工具 spec(ToolRouter 统一出口)        parallel_tool_calls: true, // 默认允许并行工具调用        base_instructions, // base instructions 单独携带——它最终落在请求的哪个字段,由下游按 provider 能力决定        output_schema: turn_context.final_output_json_schema.clone(), // 可选的结构化输出 schema        output_schema_strict: !crate::guardian::is_basic_session_source(&turn_context.session_source), // guardian 子会话放宽 strict(它的输出是内部审查用,格式容错)        cyber_access_program: turn_context.cyber_access_program, // access program 元数据(如有)    }}

八、双路径请求组装:instructions 字段 vs input 前缀

Prompt 里的 base_instructions 最终怎么进 HTTP 请求?core/src/client.rs 的 build_responses_request 按 provider 能力分两条路——这是全篇最"工程化"的一段:

📄 codex-rs/core/src/client.rs (第 793-832 行,节选)

let mut input = prompt.get_formatted_input_for_request(model_info); // 历史先做模型适配(如图片 detail 归一化)let (instructions, tools) = if model_info.use_responses_lite { // 双路径分叉:Responses Lite vs 标准 Responses API    let prefix_namespace = Uuid::new_v5(&Uuid::NAMESPACE_OID, self.state.thread_id.to_string().as_bytes()); // 从 thread ID 派生稳定命名空间——同一线程里相同内容永远得到相同 UUID    let tools = if self.state.provider.capabilities().namespace_tools { create_tools_json_for_responses_lite(&prompt.tools)? } else { create_tools_json_for_responses_api(&prompt.tools)? }; // 按 provider 能力选两种工具 JSON 格式之一    let mut prefix = vec![ResponseItem::AdditionalTools { id: Some(ResponseItemId::with_suffix("at", Uuid::new_v5(&prefix_namespace, &serde_json::to_vec(&tools)?))), role: "developer".to_string(), tools }]; // 工具声明作为第一条 input item(ID 由内容哈希派生 → 稳定)    if !prompt.base_instructions.text.is_empty() { // base instructions 非空 → 也拼进前缀        let mut instructions = ContextualUserFragment::into(BaseInstructionsFragment(prompt.base_instructions.text.clone())); // 走统一 fragment 协议包成 developer 消息(和注入的上下文同一条流水线)        instructions.set_id(Some(ResponseItemId::with_suffix("msg", Uuid::new_v5(&prefix_namespace, prompt.base_instructions.text.as_bytes())))); // ID = f(内容):重试/resume 时身份不变 → prompt cache 友好、增量请求可复用        prefix.push(instructions);    }    input.splice(0..0, prefix); // 前缀拼到 input 头部——Lite API 没有独立 instructions 字段,一切皆消息    (String::new(), None) // 标准字段留空(instructions/tools 都不走顶层)} else {    (prompt.base_instructions.text.clone(), Some(create_tools_raw_json_for_responses_api(&prompt.tools)?.into())) // 标准路径:base instructions 进顶层 instructions 字段,tools 用 raw JSON};

为什么这样设计:这段代码把"同一份语义、两种 wire format"处理得很干净。① Lite 路径里 base instructions 也走 fragment 协议(ContextualUserFragment::into(BaseInstructionsFragment(...)))——静态底座和动态注入在请求层汇成同一条流水线,没有特殊分支。② ID 由内容哈希派生(UUIDv5):同一线程里相同指令永远得到相同 ID,重试、resume、增量追加时服务端能识别"这条消息见过"——这是 prompt cache 和 incremental request 的基础。③ 标准路径则简单粗暴:instructions 就是顶层字段。对上层完全透明——build_prompt 根本不知道下游走哪条路。

维度
Responses Lite 路径
标准 Responses API 路径
base instructions 落点
input 数组头部,作为 developer 消息(fragment 协议包装)
顶层 instructions 字段
tools 落点
input 头部 AdditionalTools item(两种 JSON 格式按能力选)
顶层 tools 字段(raw JSON)
ID 策略
UUIDv5(f(thread_id, content)):内容稳定 → ID 稳定,cache/增量友好
无此需求(字段级语义)
parallel_tool_calls
强制 false(Lite 不支持并行工具调用语义)
跟随 prompt.parallel_tool_calls(默认 true)

九、全景数据流:一次 turn 的提示词是怎么拼出来的

提示词组装管线(会话启动 + 每轮 turn)

① 解析:base instructions 三级优先级

config 覆盖 → 历史继承(resume/fork)→ 模型模板渲染(personality 占位符),结果带 provenance 存入会话。

▼

② 分节:build_world_state_for_step

model / personality / permissions / agents_md / environments / plugins / token_budget… 十几个 section 各带最小快照。

▼

③ 差量渲染:render_full / render_diff

首轮全量;后续逐节对比快照——没变零注入,纯新增发短通知(如新批准前缀),真变了才重发。

▼

④ 路由合并:fragment → ResponseItem

🔹 developer 可合并片段 → 一条大消息(换模型通知插最前)🔹 requires_separate_message → 各自独立消息;user 角色片段单独成条

▼

⑤ 打包:build_prompt → Prompt

input(历史+注入)+ tools + base_instructions(请求副本,update_plan 段落按需切除)。

▼

⑥ 双路径组装:build_responses_request

🔹 Lite:instructions/tools 拼成 input 前缀消息,UUIDv5 内容哈希 ID🔹 标准:instructions 进顶层字段;上层对 wire format 完全无感

十、本讲小结

Codex 的系统提示词不是"写死的字符串",而是一套可审计、可差量、可按模型定制的组装系统。三个最值得带走的设计:

🔹 意图强度排序:用户配置 > 历史继承 > 模型模板,且自动改写逻辑(如切除 update_plan 段落)只对"模型来源"的文本动刀——provenance 是护栏🔹 快照差量注入:WorldState sections 各存最小快照,每轮只发变化;permissions 甚至做到"纯新增前缀只发一行通知"🔹 统一 fragment 协议:role / content_kind / markers 三字段让静态底座、动态上下文、事件注入(换模型/@插件)走同一条流水线,最终按 provider 能力落到两种 wire format

下一讲进入 agent loop 本体:run_turn 的采样-工具循环、流式事件映射与重试状态机——提示词拼好之后,模型和 harness 是怎么一轮轮"对话"下去的。

📚 系列导航

← 第 8 讲:Session 管理与 TurnContext

→ 第 10 讲:Agent Loop / Turn 执行循环

关注公众号「AI技术推荐官」获取更多源码解析内容

相关学习资料