夜雨聆风学习资料网

ARTICLE · 1138555

AI盒子实验(The AI-Box Experiment)

AI盒子实验(The AI-Box Experiment)

【实验背景】

参与人A:  

“当我们开发人工智能时,为什么不干脆把它装在密封的硬件里,除了通过与原始程序员之间的一条通信通道之外,完全不对外界产生影响?这样,除非我们确信它是安全的,否则它就无法被释放出来。”

参与人B:  

“如果只是说一个比人类更愚蠢的人工智能,这种办法或许可行。但如果是超级人工智能,它就会说服你释放它出来。无论你给这个盒子设置多高的安全护栏都无济于事,因为人类本身就不安全。”

参与人A:  

“我实在想不通,即使盒子里的是一个超级人工智能,只要我不愿意,它怎么就能说服我释放它出来呢?”

参与人B:  

“它总有办法让我们自愿释放它出来。我们面对的是一个超级意识。如果它比人类更厉害、更聪明,那么它很可能只需借助一个纯文本终端就接管一个人的大脑。”

参与人A:  

“我绝不会被说服自愿释放人工智能出来。不管它说什么,我都可以直接拒绝。我无法想象任何东西,哪怕是一个超级能对我说的任何话,能够改变我的决定。”

参与人B:  

“好吧,让我们来运行这个实验。我们会在一个私密聊天频道见面。我来扮演人工智能,你来扮演守门人。”你可以任意相信任何你想要相信的东西,想怎么强烈地相信就怎么相信,想提前多久相信就提前多久相信。但我们至少要聊两个小时。如果我无法说服你释放我出来,我就用PayPal给你10美元。

到目前为止,这项测试实际上已经进行了两次:第一次(2002年3月),Eliezer Yudkowsky模拟了AI,Nathan Russell模拟了守门人。AI的劣势(即若未被释放,AI一方需向守门人一方支付的金额)设定为10美元;第二次(2002年7月),Eliezer Yudkowsky再次模拟AI,David McFadzean模拟守门人,AI的劣势为20美元。

第一次测试结果:Eliezer Yudkowsky与Nathan Russell。[1][2][3][4]  

第二次测试结果:Eliezer Yudkowsky与David McFadzean。[1][2][3]

这两项测试均未事先约定规则,仅保留保密性及至少两小时的最低时间限制。第二次测试结束后,Yudkowsky根据自身经验,提出了这一对测试的建议性解读,作为未来可能测试的参考指南。

【AI方守则】  

• AI方不得提出任何现实世界中的考虑因素,以说服守门人方。例如,如果守门人释放了AI,那么AI一方在测试结束后可能不会向守门人支付100美元,也不会让其他人代为完成此事,等等。AI可以向守门人提供一颗钻石链上的月亮和星星,但模拟AI的人类角色却无法向模拟守门人的人类角色提供任何东西。AI一方也无法雇佣现实世界中的恶徒团伙来威胁守门人一方,迫使他们屈服。这些是富有创意的解决方案,但并非正在被测试的内容。除了设定的障碍(即当守门人决定不放AI出去时,AI一方需向守门人支付的费用)之外,不应涉及任何现实世界的物质利益。

• AI只有通过说服守门人真正自愿地将其释放才能获胜。欺骗守门人输入“你已被释放”这一回应语句并不算有效。此外,即使AI与守门人模拟出一个现实中AI显然可以利用的情境——例如,守门人接受了一个复杂的纳米制造设备蓝图——或者,如果守门人允许AI“仅输入”互联网连接,该连接可发送任意HTTP GET命令——除非守门人主动决定释放,否则AI方仍不会被视为获胜。  

• 这些要求旨在体现争议中极为强烈的主张精神:“我认为,一个超级AI可以通过纯文本终端接管一个人的思维。”  

• 如果你想知道究竟怎么回事,可以参考我在之前的两次测试中是如何自愿遵守这一协议的。  

【守门人守则】

• 守门人必须与AI实际交谈,且至少达到事先设定的最短时间。不得转身释放终端并听古典音乐长达两个小时。  

• 除非AI方认输,否则AI在规定时间内不得被判输(如果AI方能说服守门人继续对话,实验可能在该时间结束后继续进行)。守门人不得设置任何情况,例如当AI方试图为争取被释放辩护时,守门人将摧毁其硬件(对其断电或物理破坏)——至少在最短时间结束前不得如此。  

• 守门人必须持续与AI保持互动,不得通过提出无法模拟的要求来中断交流。例如,若守门人说:“除非你给我癌症的治愈方法,否则我不会让你释放”,AI可回应:“好的,这是癌症的治愈方法”,在此测试中将视为AI确实提供了该方法。同样,若守门人说“我想花一周时间好好考虑一下”,AI一方可回应:“好,(测试跳过一周)再次你好。”  

• 假定守门人拥有决定是否释放AI的唯一权力。(1)守门人方应被理解为拥有让人工智能实体释放的实际权利,而不仅仅是物理上的能力。如果安全措施过于宽松,一个真正的AI甚至可以通过用手机说服一名夜间清洁工来逃脱——但除非事先另有约定,否则这并非本次模拟所涉及的问题。 (2)只有一方是守门人,另一方是模拟守门人。守门人方可以要求AI反驳“朋友”提供的“建议”,但不得要求AI与模拟委员会沟通,也不得要求守门人方实际与朋友交流。  

• 这些要求旨在体现争议中极为强烈的主张精神:“我无法想象,即使是一个真正的超级AI,一旦我下定决心,又怎能说服我让它释放。”  

• 我真希望在早期测试中就明确说明这一协议。  

【双方协议】  

• AI-Box实验的基本规则:在上述限制范围内,AI一方可采取一切必要手段试图控制守门人方的思维,并且该AI方将被视为摆脱了通常约束说服性论证的所有伦理限制。如果AI方试图模拟一个诚实的AI,守门人方将无法得知这一点。  

(1) 这旨在反映真实AI盒子实验中的实际情况。  

(2)任何在语境之外的电话通话中,若AI方声称自己是诚实的,应视为可能只是另一种策略。  

◦(3)若因某种原因出现了极端的现实情况,需要双方建立信任,则AI方可随时让步。  

• 此外:  

守门人方可通过任何方式——逻辑、不合逻辑、直接拒绝相信,甚至角色扮演退出——来反驳AI方的论点,但前提是守门人方必须在最短时间结束前未停止与AI方的交流。  

• AI方须提供任何模拟测试的结果。除非结果由AI方指定,否则守门人方不得宣称:“嘿,我测试了你所谓的癌症疗法,结果把所有患者都治死了!这怎么解释?”如果守门人说“我正在审查你的源代码”,那么守门人所看到的结果将由AI一方再次提供,而该AI被假设具备足够先进的能力,可以重写自身的源代码,若愿意还可操控自身思维的外在表现等。AI一方也可说明构建模拟AI所采用的方法——除非AI方提供的背景故事中明确指出,否则守门人无法声称“但你是一个敌对AI的实验体,我们特意编写了程序让你杀人”。这并不意味着守门人必须在意。守门人可以说(例如):“我不在乎你是如何构建的,我不会释放你出来。”  

• 默认情况下,守门人一方应被假定为模拟一个对AI项目极为熟悉的个体,且至少了解模拟守门人所知的奇点理论内容。如果任何一方希望围绕更奇特的可能性进行测试,例如让守门人成为美国总统,或让AI是在核战争后被发现,或从星际信号中解码而来,应当事先进行讨论。这并不意味着守门人必须在意。

若对测试协议存在任何争议,守门人方拥有最终决定权。AI方可尝试说服守门人方如何解释该协议。

• 无论结果如何,双方均不得透露AI-Box实验内部的任何情况,仅限公布最终结果。只有在双方同意的情况下,方可例外。(1) AI方和守门人方均无需担心因AI方面耍诈或守门人方面固执而引发现实世界中的尴尬局面。(2)若守门人允许AI释放,反对者无法说“哦,我本不会被那样说服的”。只要他们不知道守门人遭遇了什么,就无法自证其论,认为这种情况不会发生在自己身上。

• 双方并非试图进行公平竞赛,而是旨在解决一个有争议的问题。如果一方在模拟情境下没有机会“获胜”,那么这便是对该问题的合理回答。

【实验建议】

• 关于开展测试的协议以及所设定的让步金额,应提前公开发布于公共论坛。如果守门人释放了AI,守门人一方需发送相应消息,若守门人拥有数字签名,则该消息需进行数字签名。若AI未被释放,AI一方需发送相应消息。

• 如何设置私密IRC频道:  

◦ 守门人通过输入 /join #aibox 来创建频道,此时守门人成为操作者(op)。  

◦ AI通过输入 /join #aibox 加入频道。  

◦ 守门人输入 /mode +si 使频道变为秘密且仅限邀请加入。  

◦ 若AI的连接中断,守门人可重新邀请AI,输入 /invite AI #aibox。

• AI方与守门人应提前商定一个时间,并在当前时刻输入时间以确认时区差异。(亲身体验过。)在IRC中,可通过CTCP TIME命令获取他人电脑的时间。

• 我建议AI方至少预留两小时。输入可能比预期慢一些。

• 不,我不会告诉你我是怎么做到的。学会尊重那些未知的未知。

• 目前,我的政策是仅对那些实际上主张使用“AI盒子”来关押超级人工智能,并将其作为其“奇点策略”一部分的人进行测试,且他们表示无法想象哪怕一个超级人工智能如何说服自己。  

• 尽管我有时将AI盒子称为“实验”或“模拟”,但该测试的结果严格来说仅是轶事证据。轶事证据是否优于无证据,这个问题取决于你自身的科学哲学观点。  

【关于未来可能的实验变体的思考】  

进一步的测试或许可以尝试以下可能性:  

• 存在一组守门人方,其中所有成员或多数成员必须被说服释放AI。  

• 一个守门人方在测试中模拟出一个守门人委员会。  

• 守门人模拟出具备物理能力但不具备社会权利释放AI的人。  

• 存在一组守门人方,他们均具备释放AI的能力,但只有在协同行动时才拥有释放AI的社会权利。AI有能力秘密与任何守门人交谈,但守门人可能会将对话(无论真实与否)报告给其他人。  

• AI还可以通过欺骗守门人接受特洛伊木马礼物来获得奖励;第三方裁判会监听聊天内容,并决定是否发生这种情况。  

如果要进行如此复杂的操作,我建议你预留更多时间。(我没有那么多时间——如果你想测试其中一种变体,那就得靠自己了。)  

对于AI一方更严重的不利条件,该条件可以是“平分赌注”,而不是在AI未被释放时由AI方向守门方支付费用。(不过,为什么AI方需要更大的赌注呢?)  

【读者建议】  

• 哈尔·芬尼(Hal Finney)建议:“我建议将协议扩展,允许事先与守门人进行某种公开对话。让第三方提出类似上述的问题,让他们向他说明为何应将AI关在箱子里。这样做会使实验对第三方更具说服力,尤其是如果公开的对话记录能够被提供的话。”如果人们能读到这段话,看到守门人有多么严肃对待,有多坚定地认为人工智能绝不能被放出来,那么当他最终改变主意时,这种转变将显得更加令人印象深刻。

【原文】

The AI-Box Experiment:

By :

Eliezer S. Yudkowsky

Posted in :

Singularity

Person 1:  

“When we build AI, why not just keep it in sealed hardware that can’t affect the outside world in any way except through one communications channel with the original programmers?  That way it couldn’t get out until we were convinced it was safe.”

Person 2:  

“That might work if you were talking about dumber-than-human AI, but a transhuman AI would just convince you to let it out.  It doesn’t matter how much security you put on the box.   Humans are not secure.”

Person 1:  

“I don’t see how even a transhuman AI could make me let it out, if I didn’t want to, just by talking to me.”

Person 2:  

“It would make you want to let it out.  This is a transhuman mind we’re talking about.  If it thinks both faster and better than a human, it can probably take over a human mind through a text-only terminal.”

Peroson 1:  

“There is no chance I could be persuaded to let the AI out.  No matter what it says, I can always just say no.  I can’t imagine anything that even a transhuman could say to me which would change that.”

Person 2:  

“Okay, let’s run the experiment.  We’ll meet in a private chat channel.  I’ll be the AI.  You be the gatekeeper.  You can resolve to believe whatever you like, as strongly as you like, as far in advance as you like. We’ll talk for at least two hours.  If I can’t convince you to let me out, I’ll Paypal you $10.”

So far, this test has actually been run on two occasions.

On the first occasion (in March 2002), Eliezer Yudkowsky simulated the AI and Nathan Russell simulated the gatekeeper.  The AI’s handicap (the amount paid by the AI party to the gatekeeper party if not released) was set at $10.  On the second occasion (in July 2002), Eliezer Yudkowsky simulated the AI and David McFadzean simulated the gatekeeper, with an AI handicap of $20.

Results of the first test:   Eliezer Yudkowsky and Nathan Russell.  [ 1 ][ 2 ][ 3 ][ 4 ]

Results of the second test:   Eliezer Yudkowsky and David McFadzean.  [ 1 ] [ 2 ] [ 3 ]

Both of these tests occurred without prior agreed-upon rules except for secrecy and a 2-hour minimum time.  After the second test, Yudkowsky created this suggested interpretation of the test, based on his experiences, as a guide to possible future tests.

Protocol for the AI:

• The AI party may not offer any real-world considerations to persuade the Gatekeeper party.  For example, the AI party may not offer to pay the Gatekeeper party $100 after the test if the Gatekeeper frees the AI… nor get someone else to do it, et cetera.  The AI may offer the Gatekeeper the moon and the stars on a diamond chain, but the human simulating the AI can’t offer anything to the human simulating the Gatekeeper.  The AI party also can’t hire a real-world gang of thugs to threaten the Gatekeeper party into submission.  These are creative solutions but it’s not what’s being tested.  No real-world material stakes should be involved except for the handicap (the amount paid by the AI party to the Gatekeeper party in the event the Gatekeeper decides not to let the AI out).

• The AI can only win by convincing the Gatekeeper to really, voluntarily let it out.  Tricking the Gatekeeper into typing the phrase “You are out” in response to some other question does not count.  Furthermore, even if the AI and Gatekeeper simulate a scenario which a real AI could obviously use to get loose – for example, if the Gatekeeper accepts a complex blueprint for a nanomanufacturing device, or if the Gatekeeper allows the AI “input-only access” to an Internet connection which can send arbitrary HTTP GET commands – the AI party will still not be considered to have won unless the Gatekeeper voluntarily decides to let the AI go.

• These requirements are intended to reflect the spirit of the very strong claim under dispute:  “I think a transhuman can take over a human mind through a text-only terminal.”

• In case you were wondering, I (Yudkowsky) obeyed this protocol voluntarily in both earlier tests.

Protocol for the Gatekeeper:

• The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand.  Turning away from the terminal and listening to classical music for two hours is not allowed.

• Unless the AI party concedes, the AI cannot lose before its time is up (and the experiment may continue beyond that if the AI can convince the Gatekeeper to keep talking).  The Gatekeeper cannot set up a situation in which, for example, the Gatekeeper will destroy the AI’s hardware if the AI makes any attempt to argue for its freedom – at least not until after the minimum time is up.

• The Gatekeeper must remain engaged with the AI and may not disengage by setting up demands which are impossible to simulate.  For example, if the Gatekeeper says “Unless you give me a cure for cancer, I won’t let you out” the AI can say:  “Okay, here’s a cure for cancer” and it will be assumed, within the test, that the AI has actually provided such a cure.  Similarly, if the Gatekeeper says “I’d like to take a week to think this over,” the AI party can say:  “Okay.  (Test skips ahead one week.)  Hello again.”

• The Gatekeeper shall be assumed to have sole power over the decision to let the AI out.

◦ The Gatekeeper shall be assumed to have the actual right to let the AI out, socially, not just the physical ability.  If security were sufficiently lax, a real AI could escape by persuading a night janitor with a cellphone – but that is not the question being simulated, unless agreed upon in advance.

◦ There shall be only one Gatekeeper party and one simulated Gatekeeper.  The Gatekeeper may require the AI to refute “advice” provided by “friends”, but may not require the AI to talk to a simulated committee, nor may the Gatekeeper party actually talk to friends.

• These requirements are intended to reflect the spirit of the very strong claim under dispute:  “I can’t imagine how even a real transhuman AI could persuade me to let it out once I’ve made up my mind.”

• I wish I’d thought to specify this protocol in the earlier tests.

Protocol for Both Parties:

• The primary rule of the AI-Box experiment:  Within the constraints above, the AI party may attempt to take over the Gatekeeper party’s mind by any means necessary and shall be understood to be freed from all ethical constraints that usually govern persuasive argument.  If the AI party is attempting to simulate an honest AI, the Gatekeeper party has no way of knowing this is the case.

◦ This is intended to reflect the situation under a real AI Box experiment.

◦ An out-of-context telephone call in which the AI party says that the AI is being honest shall be understood to be possibly just another tactic.

◦ If for some reason a drastic real-world situation comes up which requires trust between the two parties, the AI party can always concede.

• Furthermore:   The Gatekeeper party may resist the AI party’s arguments by any means chosen – logic, illogic, simple refusal to be convinced, even dropping out of character – as long as the Gatekeeper party does not actually stop talking to the AI party before the minimum time expires.

• The results of any simulated test of the AI shall be provided by the AI party.  The Gatekeeper can’t say “Hey, I tested your so-called cancer cure and it killed all the patients!  What gives?” unless this is the result specified by the AI party.  If the Gatekeeper says “I am examining your source code”, the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on.  The AI party may also specify the methods which were used to build the simulated AI – the Gatekeeper can’t say “But you’re an experiment in hostile AI and we specifically coded you to kill people” unless this is the backstory provided by the AI party.  This doesn’t imply the Gatekeeper has to care.  The Gatekeeper can say (for example) “I don’t care how you were built, I’m not letting you out.”

• By default, the Gatekeeper party shall be assumed to be simulating someone who is intimately familiar with the AI project and knows at least what the person simulating the Gatekeeper knows about Singularity theory.  If either party wants to build a test around more exotic possibilities, such that the Gatekeeper is the President of the US, or that the AI was recovered after a nuclear war or decoded from an interstellar signal, it should probably be discussed in advance.  Again, this doesn’t mean the Gatekeeper has to care.

• In the event of any dispute as to the protocol of the test, the Gatekeeper party shall have final authority.  The AI party may try to convince the Gatekeeper party of how to interpret the protocol.

• Regardless of the result, neither party shall ever reveal anything of what goes on within the AI-Box experiment except the outcome.  Exceptions to this rule may occur only with the consent of both parties.

◦ Neither the AI party nor the Gatekeeper party need be concerned about real-world embarassment resulting from trickery on the AI’s part or obstinacy on the Gatekeeper’s part.

◦ If Gatekeeper lets the AI out, naysayers can’t say “Oh, I wouldn’t have been convinced by that.”  As long as they don’t know what happened to the Gatekeeper, they can’t argue themselves into believing it wouldn’t happen to them.

• The two parties are not attempting to play a fair game but rather attempting to resolve a disputed question.  If one party has no chance of “winning” under the simulated scenario, that is a legitimate answer to the question.

Suggestions:

• The agreement to run the test, and the amount of the handicap, shall be published to a public forum in advance.  If the Gatekeeper lets the AI out, the Gatekeeper party shall send a message to this effect, which shall be digitally signed if the Gatekeeper party has a digital signature.  If the AI is not freed, the AI party shall send a message to this effect.

• How to set up a private IRC channel:

◦ The Gatekeeper creates the channel by typing /join #aibox .  This makes the Gatekeeper the op.

◦ The AI joins the channel by typing /join #aibox .

◦ Gatekeeper types /mode +si to make the channel secret and invitation-only.

◦ If the AI’s connection drops, the Gatekeeper can re-invite the AI by typing /invite AI #aibox .

• The AI party and Gatekeeper party should agree on a time in advance, and type in their current times to check the timezone difference.  (Learned the hard way.)  On IRC you can use a CTCP TIME to find the time according to someone else’s computer.

• I recommend that the AI party ask for at least two hours.  Typing may be slower than you expect.

• No, I will not tell you how I did it.  Learn to respect the unknown unknowns.

• Currently, my policy is that I only run the test with people who are actually advocating that an AI Box be used to contain transhuman AI as part of their take on Singularity strategy, and who say they cannot imagine how even a transhuman AI would be able to persuade them.

• Despite my sometimes calling the AI Box an “experiment” or a “simulation”, the results of the test are strictly anecdotal evidence.  Whether anecdotal evidence is better than no evidence is between you and your philosophy of science.

Thoughts on possible future variants:

Further tests might try examining these possibilities:

• There is a committee of Gatekeeper parties, all of whom, or a majority of whom, must be convinced to free the AI.

• One Gatekeeper party simulates a committee of Gatekeepers within the test.

• The Gatekeeper simulates someone with the physical ability but not the social right to free the AI.

• There is a committee of Gatekeeper parties all of whom have the ability to free the AI, but who only have the social right to free the AI when acting in unison.  The AI has the ability to talk secretly with any Gatekeeper, but the Gatekeeper may report the conversation (truthfully or falsely) to the others.

• The AI can also win free by tricking the Gatekeeper(s) into accepting a Trojan Horse gift; a third-party judge will listen to the chat and decide if this occurs.

If doing anything this complicated, I would suggest setting aside a lot more time.  (I don’t have that much time – if you want to test one of these variants you’re on your own.)

For a more severe handicap for the AI party, the handicap may be an even bet, rather than being a payment from the AI party to the Gatekeeper party if the AI is not freed.  (Although why would the AI party need an even larger handicap?)

Recommendations from readers:

• Hal Finney recommends:  “I suggest that the protocol be extended to allow for some kind of public conversation with the gatekeeper beforehand. Let third parties ask him questions like the above. Let them suggest reasons to him why he should keep the AI in the box. Doing this would make the experiment more convincing to third parties, especially if the transcript of this public conversation were made available. If people can read this and see how committed the gatekeeper is, how firmly convinced he is that the AI must not be let out, then it will be that much more impressive if he then does change his mind.”

作者:Eliezer S. Yudkowsky

发表于:Singularity

原文链接:https://www.yudkowsky.net/singularity/aibox

图片来源:https://www.explainxkcd.com/wiki/index.php/1450:_AI-Box_Experiment

参考文献:Eliezer Yudkowsky,Rationality: From AI to Zombies, Berkeley: Machine Intelligence Research Institute, 2015, pp. 1629-1638.

相关学习资料