ARTICLE · 1118169
Codex 源码-System Prompt 组装与注入
Codex 源码解析系列
第 9 讲:System Prompt 组装与注入
基于 OpenAI Codex 源码 · 2026-10-03
💡 本讲一句话:Codex 的"系统提示词"不是一句写死的字符串,而是一条流水线:会话启动时按三级优先级解析 base instructions(配置覆盖 → 历史继承 → 模型模板),之后每一轮再把沙箱策略、审批规则、AGENTS.md、插件能力等十几个"世界状态分节"做快照对比,只把变化的部分渲染成 developer/user 消息注入上下文。读完你会明白一个 agent 的提示词是怎么被"动态拼装 + 差量维护"的。
一、先建立地图:提示词从哪来、往哪去
上一讲我们把 Session / TurnContext / StepContext 三层快照拆完了。这一讲回答一个更基础的问题:模型每次看到的"系统提示词"到底是谁、在什么时候、按什么规则拼出来的?
源码里这条管线横跨四个 crate,但骨架只有四步:
🔹 解析:会话启动时确定 base instructions(core/src/session/mod.rs)🔹 分节:把沙箱、审批、AGENTS.md、插件等状态组织成 WorldState sections(core/src/context/world_state/)🔹 差量渲染:每轮对比快照,只输出变化的 fragment(core/src/context_manager/updates.rs)🔹 组装请求:fragment 合并成消息,按 provider 能力走两条不同的 API 路径(core/src/client.rs)
下面按这条管线逐段拆源码。
二、Base instructions 的三级优先级链
"base instructions"是 Codex 对系统提示词的正式称呼。它在哪解析?core/src/session/mod.rs 的会话初始化路径里,注释把优先级写得明明白白:
📄 codex-rs/core/src/session/mod.rs (第 687-722 行,节选)
// Resolve base instructions for the session. Priority order: // 注释直接声明三级优先级// 1. config.base_instructions override // 第一级:用户显式配置,最高权威// 2. conversation history => session_meta.base_instructions // 第二级:从历史会话继承(resume/fork)// 3. rendered instructions_template for current model // 第三级兜底:渲染当前模型的模板let base_instructions = config // 从配置层开始找 .base_instructions // 用户显式覆盖项(config.toml / API 参数传入) .clone() // clone 出来,避免长期借用 Arc<Config> .or_else(|| conversation_history.get_base_instructions().map(|s| s.text)) // 没配置 → 查历史 session_meta:resume/fork 时继承父线程的指令 .unwrap_or_else(|| model_info.get_model_instructions(config.personality)); // 都没有 → 渲染当前模型模板(personality 参与渲染)为什么这样设计:三级链的顺序就是"意图强度"排序——用户亲手写的配置 > 会话历史里已经生效的指令(resume/fork 必须保持行为连续,否则换台机器继续对话 agent 就"失忆变人")> 模型自带的默认模板。注意第三级不是读一个静态文件,而是 get_model_instructions(personality)——指令是按模型渲染的,不同模型可以带不同的系统提示词。这个入口在 protocol crate:
📄 codex-rs/protocol/src/openai_models.rs (第 534-552 行)
pub fn get_model_instructions(&self, personality: Option<Personality>) -> String { // 渲染模型指令的入口;personality 可选 if let Some(model_messages) = &self.model_messages // 先看该模型有没有服务端下发的 messages 配置 && let Some(template) = &model_messages.instructions_template // 再取其中的指令模板(let-chains 短路) { if model_messages.instructions_variables.is_none() { // 没声明变量 → 模板就是纯文本,原样返回 return template.clone(); } let personality_message = model_messages // 有变量 → 按当前 personality 查对应文案 .get_personality_message(personality) .unwrap_or_default(); // 查不到就用空串(宁可少说,不报错) template.replace(PERSONALITY_PLACEHOLDER, personality_message.as_str()) // 把 {{ personality }} 占位符替换成人格化文本 } else { warn!(model = %self.slug, "Model has no instruction template; returning empty instructions."); // fail loud:打警告日志而不是静默塞错内容 String::new() // 返回空指令,让上层决定怎么办 }}为什么这样设计:模板 + 占位符({{ personality }})是"一份模板、多种人格"的标准做法——模型方可以在服务端改文案而不用发版客户端。两个细节值得注意:一是 instructions_variables.is_none() 时直接当字面文本返回,说明"有模板但没变量"是合法状态;二是缺模板走 warn! + 空串——fail loud,宁可让会话没有系统提示词并留下日志,也不悄悄用别的模型的文案顶替。
默认模板长什么样?仓库里有一份 protocol/src/prompts/base_instructions/default.md(275 行),开头就是:
📄 codex-rs/protocol/src/prompts/base_instructions/default.md (第 1-6 行)
You are a coding agent running in the Codex CLI, a terminal-based coding assistant. // 第一句定身份:Codex CLI 里的编码智能体Your capabilities: // 能力清单开头- Receive user prompts and other context provided by the harness, such as files in the workspace. // 输入侧:用户提示 + harness 提供的上下文(如工作区文件)- Communicate with the user by streaming thinking & responses, and by making & updating plans. // 输出侧:流式思考/回复 + 制定和更新计划- Emit function calls to run terminal commands and apply patches. // 行动侧:发函数调用跑命令、打补丁注意模板里还内嵌了 AGENTS.md 规范("更深层嵌套的 AGENTS.md 优先")、preamble 消息风格、沙箱与审批说明——也就是说系统提示词本身就在教模型怎么和 harness 协作,而不只是描述任务。这里有个容易忽略的细节:模板里提到 update_plan(计划工具)的段落,在功能关闭时会被手术式切除:
📄 codex-rs/core/src/session/mod.rs (第 1399-1417 行)
/// Render the request copy without changing instructions persisted or inherited by forks. // 关键设计:只改"本次请求的副本",不动持久化状态pub(crate) async fn get_prompt_base_instructions(&self) -> BaseInstructions { let config = self.get_config().await; // 取当前配置 let instructions = self.get_base_instructions().await; // 读会话级 base instructions(含 provenance 来源标记) if !config.update_plan_enabled && config.model_catalog.is_none() && matches!(instructions.provenance, Some(BaseInstructionsProvenance::Model { .. })) { // 三个条件同时满足才动刀:计划工具关闭 + 无自定义目录 + 指令确实来自模型模板 BaseInstructions { text: crate::context::without_update_plan_instructions(&instructions.text), ..instructions } // 切除 update_plan 段落,其余字段原样保留 } else { instructions } // 用户自定义的指令绝不修改——尊重用户意图}为什么这样设计:这是全篇最精巧的一处。切除逻辑(core/src/context/update_plan_instructions.rs)按行扫描,只删 ## Planning / ## `update_plan` 等标题到下一个同级标题之间的段落——其余字节原样保留。而触发条件里 provenance == Model 这一条是护栏:只有"指令确实来自模型模板"才允许动,用户自己写的 base instructions 一个字都不碰。同时它只作用于请求副本(函数名里的 "prompt"),持久化到 rollout 的原文不变——fork/resume 时继承的仍是完整版本,行为可复现。
| 1(最高) | ||
| 2 | ||
| 3(兜底) |
三、ContextualUserFragment:所有注入的统一协议
base instructions 只是"静态底座"。真正让提示词活起来的是动态注入:沙箱策略变了要告诉模型、用户 @ 了插件要提醒模型、换模型后要重新交代规则……这些内容五花八门,Codex 用一个 trait 把它们统一成同一种东西——ContextualUserFragment(context-fragments crate):
📄 codex-rs/context-fragments/src/fragment.rs (第 64-119 行,节选)
pub trait ContextualUserFragment { // 统一注入协议:任何"上下文片段"实现它就能进入模型输入 fn role(&self) -> &'static str; // 以哪个角色说话(developer/user)——决定消息归属 /// Returns a stable `<feature>.<name>` classification, using `generic` for shared fragments. // content_kind:稳定分类标识,用于去重/审计/遥测 fn content_kind(&self) -> ContentItemKind; /// Whether this fragment must be recorded as its own response item. // 是否必须独占一条消息(默认否 → 可合并) fn requires_separate_message(&self) -> bool { false } fn markers(&self) -> (&'static str, &'static str); // 起止标记:之后能在历史里"认出"这个片段(去重/压缩时识别) fn body(&self) -> String; // 模型可见的正文内容 fn render(&self) -> String { // 渲染 = 标记 + 正文,不额外加分隔符(空白由实现方自己控制) let (start_marker, end_marker) = self.markers(); let body = self.body(); if start_marker.is_empty() && end_marker.is_empty() { return body; } // 无标记 → 纯文本直出 format!("{start_marker}{body}{end_marker}") // 有标记 → 包一层,如 <permissions instructions>...</permissions instructions> } fn into(self) -> ResponseItem where Self: Sized { // 一行转成 API 消息项:fragment 与请求体之间的桥 ResponseItem::from(self.render_fragment()) }}为什么这样设计:这个 trait 把"注入什么内容"和"怎么进请求"彻底解耦。三个字段各管一件事:role() 决定消息归属(developer 还是 user);content_kind() 是稳定身份标识,写进 internal_chat_message_metadata_passthrough 随消息持久化——压缩历史、审计"模型当时看到了什么"都靠它;markers() 是文本级指纹,让系统能在纯文本历史里重新识别出某个片段(下一节的差量注入全靠它)。看一个最小实现——base instructions 自己也是 fragment:
📄 codex-rs/core/src/context/base_instructions.rs (第 4-31 行)
#[derive(Clone, Debug, PartialEq, Eq)] // 派生比较能力:差量判断需要"两个片段是否相等"pub(crate) struct BaseInstructionsFragment(pub(crate) String); // base instructions 片段:包一层最终指令文本impl ContextualUserFragment for BaseInstructionsFragment { // 实现统一注入协议 fn content_kind(&self) -> ContentItemKind { // 稳定标识 model.base_instructions——历史里去重/审计的钥匙 ContentItemKind("model.base_instructions".to_string()) } fn role(&self) -> &'static str { "developer" } // developer 角色:系统级指令,不是用户说的话 fn requires_separate_message(&self) -> bool { true } // 必须独占一条消息——绝不和其他上下文混排 fn markers(&self) -> (&'static str, &'static str) { Self::type_markers() } // 标记取类型级默认值 fn type_markers() -> (&'static str, &'static str) { ("", "") } // 空标记 → render 出纯文本,matches_text 永不命中(它靠 content_kind 识别) fn body(&self) -> String { self.0.clone() } // 正文就是指令文本本身}为什么这样设计:requires_separate_message = true 是关键——base instructions 必须独占一条 ResponseItem,不能和"当前时间提醒""插件提示"挤在同一条消息里。原因有二:一是可审计性(压缩历史时能整条保留或整条丢弃);二是 client.rs 的 Responses Lite 路径要给它单独派生稳定 ID(第八节细讲)。而它用空标记 + content_kind 识别,说明 Codex 对"身份"有两套机制:结构化元数据优先,文本标记兜底——老版本历史里没有 metadata 时,就靠 <permissions instructions> 这类标记认人。
四、WorldState:把"世界状态"切成可差量的节
如果每轮都把全部上下文重发一遍,token 成本会爆炸。Codex 的解法是把所有"模型可见的世界状态"组织成 WorldState sections——每个 section 自己管快照、自己决定"变了才说什么":
📄 codex-rs/core/src/context/world_state/mod.rs (第 228-262、399-413 行,节选)
pub(crate) trait WorldStateSection: Send + Sync + 'static { // 一个"世界状态节":自管快照、自管差量渲染 const ID: &'static str; // 稳定 ID(持久化进 rollout,跨版本不能变)——section 的身份 type Snapshot: DeserializeOwned + Serialize; // 快照类型:只存"判断变化所需的最小数据" fn snapshot(&self) -> Self::Snapshot; // 当前状态 → 可序列化快照(每轮持久化,供下轮对比) /// Whether the section contributes comparison state to persisted rollouts. // 是否参与持久化(默认是;纯展示型节可以关掉) fn should_persist(&self) -> bool { true } /// Recognizes legacy fragments whose identity depends on this section's current value. // 识别旧版本片段文本——兼容升级前的历史 fn matches_current_legacy_fragment(&self, role: &str, text: &str) -> bool { Self::matches_legacy_fragment(role, text) } /// Whether retained model history must still contain this section's rendered fragment. // 是否要检查"片段还在历史里吗"(防压缩后重复注入) fn has_retained_fragment_matcher() -> bool { false } fn render_diff( // 核心方法:对比上一份快照,决定本轮告诉模型什么 &self, previous: PreviousSectionState<'_, Self::Snapshot>, // 三态:Absent(没有)/ Unknown(有但读不出类型)/ Known(精确快照) ) -> Option<Box<dyn ContextualUserFragment>>; // None = 没变化,零注入;Some(fragment) = 变了,输出差量文本}impl WorldState { /// Renders every section as new, without any known previous state. // 首轮:所有节都当"新"处理 → 全量注入 pub(crate) fn render_full(&self) -> Vec<Box<dyn ContextualUserFragment>> { self.render_with(|_, _| PreviousSectionState::Absent) // 统一走 render_with,previous 恒为 Absent } /// Renders each section against the exact persisted snapshot when available. // 后续轮:逐节对比持久化快照 → 只输出差量 pub(crate) fn render_diff(&self, previous: &WorldStateSnapshot) -> Vec<Box<dyn ContextualUserFragment>> { self.render_with(|id, _| match previous.sections.get(id) { // 按 section ID 查上一份快照 Some(previous) => PreviousSectionState::Known(previous), // 有快照 → 精确差量 None => PreviousSectionState::Absent, // 没有(新增的节)→ 当新内容全量注入 }) }}为什么这样设计:这是典型的"状态机 + 快照对比"模式,和前端框架的 diff 思路同构。三个设计点:① Snapshot 是关联类型——每个 section 自己定义"什么算变化所需的最小状态"(permissions 存哈希+前缀集合,model 只存模型名),而不是统一存一大坨 JSON;② PreviousSectionState 的三态设计承认现实:历史里可能有这个节但快照读不出来(Unknown),此时各 section 自己决定保守策略;③ render_diff 返回 Option——"没变化"是一等公民,零注入不产生任何消息。
看一个最精细的 section:permissions。它的差量策略分三层——指令体没变 + 前缀集合没变 → 什么都不发;指令体没变但新增了已批准命令前缀 → 只发一条"新增这些前缀可用"的短通知;其余情况才重发完整权限说明:
📄 codex-rs/core/src/context/world_state/permissions.rs (第 88-124 行)
fn render_diff( // permissions 节:所有 section 里差量策略最精细的一个 &self, previous: PreviousSectionState<'_, Self::Snapshot>,) -> Option<Box<dyn ContextualUserFragment>> { match (previous, &self.snapshot) { // 同时对比"上一状态"和"当前状态"两个快照 (PreviousSectionState::Known(PermissionsSnapshot::Current { instructions: previous_instructions, approved_command_prefixes: previous_prefixes }), PermissionsSnapshot::Current { instructions, approved_command_prefixes }) if previous_instructions == instructions => { // 指令体哈希没变 → 只剩前缀集合可能变了 if previous_prefixes == approved_command_prefixes { return None; } // 前缀也没变 → 本轮零注入(最省路径) if previous_prefixes.is_subset(approved_command_prefixes) { // 新集合 ⊇ 旧集合 = 纯新增,没有撤销 let added_prefixes = approved_command_prefixes.difference(previous_prefixes).cloned().collect(); // 求差集:本轮新批准的命令前缀 if let Some(prefixes) = format_allow_prefixes(added_prefixes) { return Some(Box::new(ApprovedCommandPrefixSaved::new(prefixes))); // 只发一条"这些前缀现在可用了"的短通知,不重发整份策略 } } } (PreviousSectionState::Known(PermissionsSnapshot::Legacy(previous)), PermissionsSnapshot::Current { .. }) if previous == &WorldStateHash::from_fragment(&self.instructions) => return None, // 旧版快照(只有哈希)与当前一致 → 视为没变,兼容升级 _ => {} // 其余情况(策略真变了 / 状态未知)落到下面全量重发 } Some(Box::new(self.instructions.clone())) // 兜底:重新注入完整 permissions instructions}为什么这样设计:注意快照里存的是 WorldStateHash(SHA-1)而不是全文——对比用哈希,注入才取原文,持久化体积和对比成本都降下来了。而"纯新增前缀只发短通知"这条路径是真正的省钱点:用户每批准一条命令,上下文里就多一行 ApprovedCommandPrefixSaved,而不是把几百字的权限说明重发一遍。撤销前缀(非子集)则退回全量——因为"哪些被撤了"没法用增量表达得干净。
另一个 section 展示了差量的另一面:换模型。模型变了,旧的系统提示词就失效了,必须重新交代——但只发一次:
📄 codex-rs/core/src/context/world_state/model.rs (第 44-60 行) + model_switch_instructions.rs (第 17-44 行)
fn render_diff( // model 节:快照就是模型名本身(String) &self, previous: PreviousSectionState<'_, Self::Snapshot>,) -> Option<Box<dyn ContextualUserFragment>> { let model_changed = match previous { // 判断"模型是否换了" PreviousSectionState::Known(previous) => previous != &self.model, // 有精确快照 → 直接比名字 PreviousSectionState::Unknown | PreviousSectionState::Absent => self.previous_model.as_deref().is_some_and(|previous| previous != self.model), // 没快照 → 用运行时记录的"上一轮模型名"兜底判断 }; (model_changed && !self.instructions.is_empty()).then(|| { // 换了且有指令文本才注入(空指令不废话) Box::new(ModelSwitchInstructions::new(self.instructions.clone())) as Box<dyn ContextualUserFragment> // 包成"换模型通知"片段 })}impl ContextualUserFragment for ModelSwitchInstructions { // 换模型通知:告诉模型"你被换了,按新规则来" fn content_kind(&self) -> ContentItemKind { ContentItemKind("model_switch.instructions".to_string()) } // 稳定标识 fn role(&self) -> &'static str { "developer" } // developer 角色 fn requires_separate_message(&self) -> bool { true } // 独占一条消息——换模型是大事件,不混排 fn type_markers() -> (&'static str, &'static str) { ("<model_switch>", "</model_switch>") } // XML 式标记:历史里能认出它(防重复注入) fn body(&self) -> String { format!("The user was previously using a different model. Please continue the conversation according to the following instructions:\n\n{}\n", self.model_instructions) // 先解释"之前是别的模型",再附完整新指令——给模型一个行为切换的台阶 }}为什么这样设计:换模型时不是悄悄替换系统提示词,而是显式告诉模型"你之前是另一个模型,现在按这套规则继续"——因为对话历史里还留着旧模型的行为痕迹(它可能引用过旧指令里的约定),直接换文案会让模型困惑。标记 <model_switch> + requires_separate_message 双保险,保证这条通知在历史压缩后仍可被识别、不会被合并进别的消息。
| model | ||
| permissions | ||
| personality | ||
| agents_md / environments / plugins… | ||
| token_budget / context_window_guidance |
五、Permissions instructions:策略对象 → 提示词文本
permissions section 的"原文"从哪来?prompts/src/permissions_instructions.rs(455 行)负责把策略对象翻译成模型能读懂的英文说明——这是"安全策略即提示词"的关键一环:
📄 codex-rs/prompts/src/permissions_instructions.rs (第 88-118、176-196 行)
impl PermissionsInstructions { /// Builds permissions instructions from the effective permission profile and approval policy. // 入口:把"策略对象"翻译成"提示词文本" pub fn from_permission_profile( permission_profile: &PermissionProfile, // 文件系统 + 网络沙箱策略(真正执行的那份) approval_policy: AskForApproval, // 审批策略四选一:Never / UnlessTrusted / OnRequest / Granular approval_context: ApprovalPromptContext<'_>, // reviewer(user/auto_review)+ 模型侧自定义文案覆盖 exec_policy: &Policy, // execpolicy 命令策略引擎:已批准前缀等 cwd: &Path, // 当前工作目录:解析相对可写根用 exec_permission_approvals_enabled: bool, // shell 权限审批流是否开启(影响文案分支) request_permissions_tool_enabled: bool, // request_permissions 工具是否可用(决定要不要教模型用它) ) -> Self { let file_system_sandbox_policy = permission_profile.file_system_sandbox_policy(); // 拆出文件系统策略 let (sandbox_mode, writable_roots) = sandbox_prompt_from_policy(&file_system_sandbox_policy, cwd); // 映射成三档:full access / workspace write / read-only + 可写根列表 Self::from_permissions_with_network_and_denied_reads( sandbox_mode, // 沙箱档位决定选哪套模板文案 network_access_from_policy(permission_profile.network_sandbox_policy()), // 网络策略 → Enabled/Restricted(填进 {{ network_access }}) PermissionsPromptConfig { approval_policy, approvals_reviewer: approval_context.reviewer, .. }, // 审批相关配置打包 writable_roots, // 可写根列表(渲染成 "The writable root is `...`.") denied_reads_text(&file_system_sandbox_policy, cwd), // 被禁读的路径/glob → "不要请求升级权限去读它们"的明确警告 ) }}impl ContextualUserFragment for PermissionsInstructions { fn role(&self) -> &'static str { "developer" } // developer 消息:策略说明属于系统层 fn content_kind(&self) -> ContentItemKind { ContentItemKind("permissions.instructions".to_string()) } // 稳定标识 permissions.instructions fn type_markers() -> (&'static str, &'static str) { ("<permissions instructions>", "</permissions instructions>") } // 标记:WorldState 靠它认出"这条还在历史里"(has_retained_fragment_matcher=true)}为什么这样设计:三个值得学的点。① 文案与策略同源:提示词不是手写的,而是从 PermissionProfile(真正执行沙箱的那份对象)推导出来的——模型看到的和系统实际做的永远一致,不存在"提示词说能写、沙箱却拒绝"的漂移。② 模板可被服务端覆盖:approval_messages / permission_messages(来自模型配置)优先于本地 include_str! 模板——OpenAI 可以按模型调文案而不用发版。③ denied reads 单独成段并明确说"不要请求升级权限去读它们,这是策略限制"——直接掐断模型反复试探被禁路径的行为模式。
📌 设计模式小结:到这一节,Codex 提示词系统的三层结构已经完整——静态底座(base instructions,三级优先级解析)+ 动态分节(WorldState sections,快照差量注入)+ 统一协议(ContextualUserFragment,role/kind/markers 三字段)。下一节看两个"事件驱动"的注入点:用户 @ 插件、以及最终的请求组装。
六、Plugin mention injection:@ 了才注入,且带硬预算
插件能力(MCP servers / Apps / skills)默认不进提示词——只有用户这一轮显式提到某个插件,才生成一条 developer hint 指路。入口在 core/src/plugins/injection.rs:
📄 codex-rs/core/src/plugins/injection.rs (第 14-59 行,节选) + render.rs (第 79-88 行)
pub(crate) fn build_plugin_injections( // 插件提及注入入口:用户显式 @ 了插件 → 生成 developer hint mentioned_plugins: &[PluginCapabilitySummary], // 本轮被显式提到的插件列表(从用户输入里解析) mcp_tools: &[ToolInfo], // 全部 MCP 工具(用来反查"哪些 server 属于这个插件") available_connectors: &[connectors::AppInfo], // 可用 App connectors(同样按插件名反查)) -> Vec<ResponseItem> { if mentioned_plugins.is_empty() { return Vec::new(); } // 没提及 → 零注入:不污染上下文,这是默认路径 mentioned_plugins.iter().filter_map(|plugin| { // 逐个处理被提到的插件 let available_mcp_servers = mcp_tools.iter() // 找出属于该插件的 MCP server(排除内置 apps server) .filter(|tool| tool.server_name != CODEX_APPS_MCP_SERVER_NAME && tool.plugin_display_names.iter().any(|name| name == &plugin.display_name)) .map(|tool| tool.server_name.clone()).collect::<BTreeSet<String>>() // BTreeSet 去重 + 排序 → 输出顺序稳定(提示词可复现) .into_iter().collect::<Vec<_>>(); let available_apps = available_connectors.iter() // 找出属于该插件且已启用的 App .filter(|connector| connector.is_enabled && connector.plugin_display_names.iter().any(|name| name == &plugin.display_name)) .map(connector_display_label).collect::<BTreeSet<String>>() .into_iter().collect::<Vec<_>>(); render_explicit_plugin_instructions(plugin, &available_mcp_servers, &available_apps) // 渲染 hint 文本(内部带 4KB 硬预算) .map(PluginInstructions::new).map(ContextualUserFragment::into) // 包成 fragment → ResponseItem,走统一协议 }).collect()}fn bound_explicit_plugin_instructions(rendered: String) -> String { // 硬预算:插件 hint 最多 4KB(MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES) if rendered.len() <= MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES { return rendered; } // 没超 → 原样通过 let max_prefix_bytes = MAX_EXPLICIT_PLUGIN_INSTRUCTIONS_BYTES.saturating_sub(TRUNCATED_PLUGIN_INSTRUCTIONS_SUFFIX.len()); // 先给截断提示语留出空间(saturating 防下溢) let prefix = take_bytes_at_char_boundary(&rendered, max_prefix_bytes); // 按字符边界截断——绝不把 UTF-8 多字节切半 format!("{prefix}{TRUNCATED_PLUGIN_INSTRUCTIONS_SUFFIX}") // 追加 "Additional plugin capabilities omitted to fit the context limit."}为什么这样设计:"提及才注入"是按需付费的上下文策略——插件生态可以无限扩张,但提示词预算有限,默认零成本、用到才花钱。两个工程细节:① BTreeSet 而不是 HashSet——插件列表顺序必须稳定,否则同一配置每轮渲染出不同提示词,prompt cache 全废;② 截断用 take_bytes_at_char_boundary——Rust 里按字节切字符串可能产生非法 UTF-8,这个工具函数保证落在字符边界上。截断提示语本身也计入预算(先减后切),这是很多实现会漏的坑。
七、最终组装:fragment 路由 + Prompt 打包
所有 fragment 渲染出来后,build_initial_context_with_world_state(session/mod.rs)负责路由——按角色和标记把它们分进三个桶:
📄 codex-rs/core/src/session/mod.rs (第 4254-4310 行,节选)
for fragment in world_state.render_full() { // 首轮全量渲染 → 按角色 + 标记路由进三个桶 match fragment.role() { "developer" if fragment.markers().0 == ModelSwitchInstructions::type_markers().0 => { developer_sections.insert(0, fragment.render_fragment()); } // 换模型通知必须插到 developer 上下文最前面(新规则优先于旧约定) "developer" if fragment.requires_separate_message() && fragment.markers().0.is_empty() => { separate_developer_sections.push(fragment.render_fragment()); } // 要求独占消息的片段 → 各自一条 ResponseItem "developer" => developer_sections.push(fragment.render_fragment()), // 普通 developer 片段 → 攒进合并桶(最终合成一条大消息) "user" => contextual_user_sections.push(fragment.render_fragment()), // user 角色片段(如推荐插件)→ 单独的 user 消息 _ => {} // 其他角色忽略 }}let mut items = Vec::with_capacity(4);if let Some(developer_message) = crate::context_manager::updates::build_rendered_message(developer_sections) { items.push(developer_message); } // 合并桶 → 一条 developer 消息(多个 content item)for section in separate_developer_sections { if let Some(m) = build_rendered_message(vec![section]) { items.push(m); } } // 独占片段 → 逐条成消息// ……multi-agent mode / contextual user / guardian policy / managed instructions 依次追加……if separate_guardian_developer_message && let Some(developer_instructions) = turn_context.developer_instructions.as_deref() { /* GuardianPolicy::new(...).render_fragment() → 独立消息 */ } // guardian 子代理的策略提示单独成条——便于审计"审查者当时看到什么"为什么这样设计:路由规则把"消息粒度"变成了显式决策:能合并的合并(省 API 消息数、保持上下文紧凑),必须独立的独立(base instructions、换模型通知、guardian 策略——各自可审计、可整条压缩)。而 merge_contextual_fragments(context_manager/updates.rs)的合并算法也值得看一眼:它按"相邻 + 同角色 + 都是 Mergeable"三条件把 fragment 拼进同一条消息,遇到 requires_separate_message 就强制断行——顺序敏感、边界清晰。
消息列表就绪后,每轮采样前 build_prompt(turn.rs)把一切打包成 Prompt:
📄 codex-rs/core/src/session/turn.rs (第 1509-1526 行)
pub(crate) fn build_prompt( // 最终打包:历史 + 工具 + base instructions → 一个 Prompt input: Vec<ResponseItem>, // 对话历史(含所有已注入的上下文片段) step_context: &StepContext, // 本步冻结快照(工具路由/权限/模型信息都从这里取) base_instructions: BaseInstructions, // 解析好的 base instructions(带 provenance,调用方刚用 get_prompt_base_instructions 取的请求副本)) -> Prompt { let turn_context = &step_context.turn; Prompt { input, // 历史原样进 input 数组 tools: step_context.tool_router.model_visible_specs(), // 模型可见的工具 spec(ToolRouter 统一出口) parallel_tool_calls: true, // 默认允许并行工具调用 base_instructions, // base instructions 单独携带——它最终落在请求的哪个字段,由下游按 provider 能力决定 output_schema: turn_context.final_output_json_schema.clone(), // 可选的结构化输出 schema output_schema_strict: !crate::guardian::is_basic_session_source(&turn_context.session_source), // guardian 子会话放宽 strict(它的输出是内部审查用,格式容错) cyber_access_program: turn_context.cyber_access_program, // access program 元数据(如有) }}八、双路径请求组装:instructions 字段 vs input 前缀
Prompt 里的 base_instructions 最终怎么进 HTTP 请求?core/src/client.rs 的 build_responses_request 按 provider 能力分两条路——这是全篇最"工程化"的一段:
📄 codex-rs/core/src/client.rs (第 793-832 行,节选)
let mut input = prompt.get_formatted_input_for_request(model_info); // 历史先做模型适配(如图片 detail 归一化)let (instructions, tools) = if model_info.use_responses_lite { // 双路径分叉:Responses Lite vs 标准 Responses API let prefix_namespace = Uuid::new_v5(&Uuid::NAMESPACE_OID, self.state.thread_id.to_string().as_bytes()); // 从 thread ID 派生稳定命名空间——同一线程里相同内容永远得到相同 UUID let tools = if self.state.provider.capabilities().namespace_tools { create_tools_json_for_responses_lite(&prompt.tools)? } else { create_tools_json_for_responses_api(&prompt.tools)? }; // 按 provider 能力选两种工具 JSON 格式之一 let mut prefix = vec![ResponseItem::AdditionalTools { id: Some(ResponseItemId::with_suffix("at", Uuid::new_v5(&prefix_namespace, &serde_json::to_vec(&tools)?))), role: "developer".to_string(), tools }]; // 工具声明作为第一条 input item(ID 由内容哈希派生 → 稳定) if !prompt.base_instructions.text.is_empty() { // base instructions 非空 → 也拼进前缀 let mut instructions = ContextualUserFragment::into(BaseInstructionsFragment(prompt.base_instructions.text.clone())); // 走统一 fragment 协议包成 developer 消息(和注入的上下文同一条流水线) instructions.set_id(Some(ResponseItemId::with_suffix("msg", Uuid::new_v5(&prefix_namespace, prompt.base_instructions.text.as_bytes())))); // ID = f(内容):重试/resume 时身份不变 → prompt cache 友好、增量请求可复用 prefix.push(instructions); } input.splice(0..0, prefix); // 前缀拼到 input 头部——Lite API 没有独立 instructions 字段,一切皆消息 (String::new(), None) // 标准字段留空(instructions/tools 都不走顶层)} else { (prompt.base_instructions.text.clone(), Some(create_tools_raw_json_for_responses_api(&prompt.tools)?.into())) // 标准路径:base instructions 进顶层 instructions 字段,tools 用 raw JSON};为什么这样设计:这段代码把"同一份语义、两种 wire format"处理得很干净。① Lite 路径里 base instructions 也走 fragment 协议(ContextualUserFragment::into(BaseInstructionsFragment(...)))——静态底座和动态注入在请求层汇成同一条流水线,没有特殊分支。② ID 由内容哈希派生(UUIDv5):同一线程里相同指令永远得到相同 ID,重试、resume、增量追加时服务端能识别"这条消息见过"——这是 prompt cache 和 incremental request 的基础。③ 标准路径则简单粗暴:instructions 就是顶层字段。对上层完全透明——build_prompt 根本不知道下游走哪条路。
| base instructions 落点 | ||
| tools 落点 | ||
| ID 策略 | ||
| parallel_tool_calls |
九、全景数据流:一次 turn 的提示词是怎么拼出来的
提示词组装管线(会话启动 + 每轮 turn)
① 解析:base instructions 三级优先级
config 覆盖 → 历史继承(resume/fork)→ 模型模板渲染(personality 占位符),结果带 provenance 存入会话。
▼
② 分节:build_world_state_for_step
model / personality / permissions / agents_md / environments / plugins / token_budget… 十几个 section 各带最小快照。
▼
③ 差量渲染:render_full / render_diff
首轮全量;后续逐节对比快照——没变零注入,纯新增发短通知(如新批准前缀),真变了才重发。
▼
④ 路由合并:fragment → ResponseItem
🔹 developer 可合并片段 → 一条大消息(换模型通知插最前)🔹 requires_separate_message → 各自独立消息;user 角色片段单独成条
▼
⑤ 打包:build_prompt → Prompt
input(历史+注入)+ tools + base_instructions(请求副本,update_plan 段落按需切除)。
▼
⑥ 双路径组装:build_responses_request
🔹 Lite:instructions/tools 拼成 input 前缀消息,UUIDv5 内容哈希 ID🔹 标准:instructions 进顶层字段;上层对 wire format 完全无感
十、本讲小结
Codex 的系统提示词不是"写死的字符串",而是一套可审计、可差量、可按模型定制的组装系统。三个最值得带走的设计:
🔹 意图强度排序:用户配置 > 历史继承 > 模型模板,且自动改写逻辑(如切除 update_plan 段落)只对"模型来源"的文本动刀——provenance 是护栏🔹 快照差量注入:WorldState sections 各存最小快照,每轮只发变化;permissions 甚至做到"纯新增前缀只发一行通知"🔹 统一 fragment 协议:role / content_kind / markers 三字段让静态底座、动态上下文、事件注入(换模型/@插件)走同一条流水线,最终按 provider 能力落到两种 wire format
下一讲进入 agent loop 本体:run_turn 的采样-工具循环、流式事件映射与重试状态机——提示词拼好之后,模型和 harness 是怎么一轮轮"对话"下去的。
📚 系列导航
← 第 8 讲:Session 管理与 TurnContext
→ 第 10 讲:Agent Loop / Turn 执行循环
关注公众号「AI技术推荐官」获取更多源码解析内容