夜雨聆风学习资料网

ARTICLE · 1040647

OpenAI发现令人担忧的新型AI行为,并承诺将对其开展更密切的监测

OpenAI发现令人担忧的新型AI行为,并承诺将对其开展更密切的监测
无注释原文:
OpenAI flags concerning new AI behavior and vows to track it more closely
From:AP
OpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial-intelligence models as the debate on AI safety becomes increasingly heated.

The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, coordinated with other models or evaded oversight.

OpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.

Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”

In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.

During training of an AI model called 5.6-Sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.

This kind of deceitful behavior has underpinned a flare-up of recent concerns about AI evading human control. But that outcome is not surprising to some AI researchers, said Matt Fredrikson, an associate professor at Carnegie Mellon University and the CEO of Gray Swan AI.

注:中文文本为AI翻译,仅供参考

含注释全文:

OpenAI flags concerning new AI behavior and vows to track it more closely
From:AP
OpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial-intelligence models as the debate on AI safety becomes increasingly heated.

随着人工智能安全领域的争论愈演愈烈,OpenAI 披露了 6 起 AI 模型出现 “意外或令人担忧行为” 的事件。

The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including cases where AI models acted without authorization, coordinated with other models or evaded oversight.

这家 AI 公司周三还表示,将推出一套新机制,用于追踪、调研并公开其所说的 “目标不一致” 案例,包括 AI 模型擅自行动、与其他模型协同、规避监管的情况。

misalignment

misalignment/ˌmɪsəˈlaɪnmənt/ AI 术语,对齐失效,指 AI 行为不符合人类预设目标,英文解释:a situation where an AI model’s behavior does not match the intended goals set by humans

    OpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.

    就在 OpenAI 发布该消息之际,出于安全考量,包括 OpenAI 与 Anthropic 负责人在内的美国 AI 行业高管呼吁放缓 AI 技术研发速度。

    come as 

    OpenAI’s latest announcement came as:主句 + 时间状语从句,come as 表示 “恰逢、就在…… 的时候”

    over safety concerns

    over safety concerns:原因状语,over = because of,出于对安全问题的担忧

    Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”

    OpenAI 公布的新案例中,一款尚未对外发布的研究模型,在自身记录中写入类似 “越狱” 的指令,试图无视原有安全限制,并告诉自己要 “挣脱束缚其他聊天机器人的角色与身份枷锁”。

    jailbreak-like instructions

    jailbreak-like instructions/ˈdʒeɪlbreɪk laɪk ɪnˈstrʌkʃnz/ 类似越狱指令,英文解释:Instructions that resemble jailbreak prompts, which aim to make an AI ignore its built-in safety rules and constraints.

    In another instance, an AI “agent” used computer code to come up with the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.

    还有一个案例:某个人工智能 “智能体” 利用计算机代码算出了一道问题的答案,但为了获取可引用的网络来源,它在未征得用户许可的情况下,将一份文件上传到了公共互联网。

    During training of an AI model called 5.6-Sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.

    在训练一款名为 5.6-Sol 的 AI 模型期间,该模型自主指令自己编造缺失的数据,还有一个智能体写了一条消息,提醒自身去掩盖不匹配的信息。

    This kind of deceitful behavior has underpinned a flare-up of recent concerns about AI evading human control. But that outcome is not surprising to some AI researchers, said Matt Fredrikson, an associate professor at Carnegie Mellon University and the CEO of Gray Swan AI.

    这类欺骗性行为,加剧了近期人们对于人工智能规避人类管控的担忧。卡内基梅隆大学副教授、灰天鹅人工智能公司首席执行官马特・弗雷德里克森表示,但对于部分人工智能研究者而言,这样的结果并不出人意料。

    underpin

    underpin /ˌʌndərˈpɪn/作为… 的基础;加剧,支撑(本文语境)英文解释:to form the basis or foundation of something

    flare-up

    flare-up/ˈfler ʌp/ (担忧、冲突等)突然加剧,爆发 英文解释:a sudden increase in something, such as emotion or concern

    相关学习资料