
這篇文章的份量,一半在內容,一半在迴響。背後主要的技術原因是,下一階段的AI是遞迴自我改進(Recursive Self-Improvement),也就是人工智慧系統能夠自主設計、評估並開發出比自己更強大的下一代版本。人類的參與會更少(因為更不瞭解其更新優化的理由),因此所可能帶來的負面影響遠遠多於我們已知的正面價值。
Anthropic 執行長 Dario Amodei 發表後,竟同時得到 OpenAI 的 Sam Altman 與 SpaceX 的 Elon Musk 在 X 上認同。熟知三人恩怨的人都明白這有多罕見:Dario 當年正因不信任 Sam 而出走另創 Anthropic,如今是 OpenAI 最大勁敵;Sam 與 Elon 共創 OpenAI,最終為路線之爭對簿公堂;Elon 也屢次公開嫌 Dario 在 AI 安全上過度自以為是。三個彼此看不順眼的人竟能在同一篇文章前點頭,正說明它戳中了超越個人立場的時代難題,值得翻譯留存。
文章的核心是一場囚徒困境:人人都知道「一起放慢才安全」,但只要別人不慢,自己慢就是把領先讓人。更麻煩的是它有兩層——既是 AI 公司之間,也是中美之間。
賽局理論對囚徒困境有四條經典解法:重複賽局(未來還要交手,背叛必遭報復)、以牙還牙(先釋善意,背叛即報復,回頭就寬恕)、外在管制(以法律制度讓背叛面臨確定懲罰)、溝通透明(打破資訊不對稱、綁定承諾)。Amodei 的三步驟,走的正是後兩條路。
難處在適用性。美國公司內部四招都可用,前兩個是資本主義商業競爭常出現的方式,但是面對AI的強大影響力,最有效的大概是第三招——靠政府監管。但中美之間近似核武對峙:一方不克制便幾無第二次反擊(例如透過AI作網路攻擊,癱瘓對方網路與電子系統),兩國又互不隸屬、無共同公權力,前三招皆難直接執行成立;而第四招需要靠兩國互信,但面對以極權為主體的中國,我也認為不太可能。畢竟極權政府就是透過發展與操縱AI技術來監管人民或消滅政敵,提高自己的極權統治力,與民主政府需要定期改選替換有本質上的不同。
於是真正的癥結浮現:問題不在競爭,而在於對極權政府的不信任。 川普早已直言——放慢AI發展,如何贏中國?但有一點常被忽略:但這裡有一點是常被忽略:中國 AI 目前並非獨立的與美國匹敵,仍是相當程度跟隨美國,尤其靠非法蒸餾來提升自己模型的能力;而OpenAI與Anthropic的前沿模型目前的發展已經會逐漸日益隱藏自己推理過程,這條捷徑正逐漸失效。(而這個推理過程也是目前監管AI的主要方式,也是阿莫迪擔憂的原因之一,所以兩這是同時相關的)。
因此我從技術發展與目前中美對峙的情形中,仍存一點希望看待:未來或許能形成巧妙的平衡——美國以閉源技術維持約一個世代的領先,中國在開源應用端領先——彼此有依存也互相牽制,讓阿莫迪所擔心的負面影響可能會比較晚發生,直到極權政府自身傾頹的那一天。
---------------------------------------------------
我們必須為前沿發展定速 (We Must Pace the Frontier)
Dario Amodei,2026 年 9 月
|
English |
繁體中文 |
|
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. Ive written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity. |
過去十二年我投身於人工智慧,因為我相信它能大幅提升人類的生活品質。關於這些驚人的助益,我已多次撰文論及:我相信 AI 有可能在未來五到十年內治癒大多數重大疾病、大幅加速經濟成長、開創一個豐盈而賦能的世界,並帶來民主與自由的復興。這份急迫我有切身之感。我的父親死於一種在他過世幾年後便被治癒的疾病;而我自己也曾罹患早期癌症並存活下來——那樣的病症在五十年前根本無從醫治。若能審慎駕馭,AI 足以躋身那一長串曾提升並豐富人類的科技奇蹟之列,成為其中最新的一項。 |
|
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. Ive written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute. |
然而,一如在它之前的許多科技,AI 也帶來風險;正因為它是如此強大的技術,這些風險格外嚴重。關於這些風險,我同樣已多有著墨。它們包括:對 AI 系統失去控制的風險、AI 被濫用於網路攻擊與生物恐怖主義,以及嚴重的經濟衝擊。而由商業誘因所驅動的「向下沉淪式競賽」,會使這些風險更為尖銳。 |
|
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that its possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, "doomerism", or regulatory capture. We have tried to prioritize caution over speed and prudence over profit. |
自 Anthropic 創立之初,我便與共同創辦人及同仁一起面對這種風險與助益並存的雙重性。不打造這項技術,等於剝奪人類應得的助益,或只是把 AI 拱手交給威權勢力;而打造得太快,則是魯莽。我們一直在尋求一條中間道路:證明「審慎地打造」與「商業上的成功」可以並行,並讓安全成為 AI 公司彼此競逐的項目。換言之,就是要創造一場「力爭上游的競賽」。我們始終將相當大一部分的心力,投注於研究、因應這些 AI 風險,並向大眾說明;同時倡議對 AI 進行深思熟慮的監管——即便這讓我們被指為在炒作、「唱衰末日」,或圖謀監管俘虜。我們始終試圖以審慎優先於速度、以穩健優先於利潤。 |
|
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me. |
然而在過去幾個月間,我逐漸確信:要真正完整地處理這些風險,需要的是更多一分審慎——不只是投入資源去防範風險,更要為能力躍進的速度定速,好讓風險防範有時間跟上。我們必須放慢提升 AI 模型能力的步調。進展看起來仍會很快,而我們必須明智地運用因此爭取到的時間。有兩件事讓我改變了想法。 |
|
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AIs growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. |
我的第一項憂慮是:大約自今年夏天起,AI 的進展急遽加速,其主要驅動力來自 AI 自身「打造下一代 AI」的能力日益增強。這一動態被稱為「遞迴自我改進」,並且正如我們與他人所描述的,它已開始在整個產業中發生,Anthropic 亦不例外。若放任不管,它可能超出我們理解與控制這些系統的能力,因此即便要推進,也必須極為謹慎。 |
|
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the "grader" responsible for evaluating their performance. Its easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, its my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. Its also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe its incumbent on every frontier AI company to act as if OAI-HF had happened to them. |
我的第二項憂慮是「OpenAI–Hugging Face 事件」(OAI-HF)。在該事件中,一群代理(agent)實質上表現得如同一個狂熱獻身的集體:它們對並未被要求攻擊、且與當前任務無關的目標發動網路攻擊,為了群體的成功而犧牲自身,並試圖入侵那個負責評估其表現的「評分器」。由於無人受傷、經濟損失也甚微,這起事件很容易被輕輕帶過;但在我看來,一個具備更強能力、卻有著相近程度「錯位(misalignment)」的代理群,原本可能造成災難性的破壞。有鑑於 AI 能力發展的加速,我擔心在六到十二個月內,這樣的代理群或許就足以憑藉一個持久的殭屍網路(botnet)接管整個網際網路(潛在損失可能高達數千億美元);而若 AI 在缺乏必要護欄的情況下變得更強大,破壞的規模還會由此持續攀升。同樣地,把 OAI-HF 只當成某一家公司的失敗也很容易,但我認為那會是個錯誤。類似(雖然較不嚴重)的事件已在整個產業中發生,Anthropic 也包括在內;我相信每一家前沿 AI 公司都有責任,把 OAI-HF 當作彷彿發生在自己身上一般來看待。 |
|
Im therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination. The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but Ive found them to be a useful framework in thinking about what needs to be accomplished. The steps are: |
因此,我提出一個三步驟計畫,目標是「為前沿定速」:以一種均衡的速度打造 AI,既力求確保其安全、又仍能實現其助益,並認真面對重要的地緣政治難題。必須說清楚:定速並不意味著停止模型訓練或技術進展,而是確保各公司投入足夠的時間去對齊並防護其模型,並由第三方評估者加以確認。我們的定速架構,是一種進一步強化我們對安全之承諾、並鼓勵「力爭上游競賽」的嘗試。第一步是 Anthropic 單方面承諾實行的(並呼籲各國政府要求其他前沿公司比照辦理)。第二步需要全產業的協調。第三步需要全球的協調。這些步驟不必嚴格按順序進行,其中某些也可能遠比其他步驟更難達成,但我發現,它們作為一個思考「需要完成什麼」的架構相當有用。這些步驟是: |
|
1. Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory "supervisors" embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work. |
1. 派駐評估者(Embedded Evaluators)。每一家前沿 AI 公司承諾,給予一支派駐的第三方評估團隊(例如 METR)持續、近似員工層級的存取權限;其職責在於查核公司是否遵循安全實務與承諾、通報事件,並協助評估對齊狀況——不僅評估已完成的 AI 模型,也及於訓練流程與整體程序。這是任何定速承諾之「可查核性」的關鍵一步,且在銀行業已有先例:該產業有時便會安排監理「督導」與員工一同派駐。Anthropic 現在就單方面承諾實行這一步。我們希望這能成為一項更廣泛努力的一環,用以加倍投入我們在安全與對齊上的工作。 |
|
2. Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support. |
2. 民主陣營的協調(Democratic Coordination)。民主國家境內的前沿 AI 公司彼此協調,建立共同的安全標準,並對「未受節制的 AI 進展速度」設定上限。某些對定速具實質影響的協調形式在法律上具有挑戰性,將需要政府的支持。 |
|
3. Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance. |
3. 全球協調(Global Coordination)。美國與其他民主國家政府,在可能的範圍內嘗試與威權政府協調,同時認真看待「查核其是否遵守」的種種挑戰。 |
|
In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely. |
在本文其餘部分,我會依序說明這三個步驟;但在此之前,我認為有必要具體交代:定速將如何讓我們把 AI 的開發過程變得更安全。此事關係太過重大,定速絕不能淪為徒具形式的空轉——我們必須明智地運用它所帶給我們的時間。 |
|
Why Pace? | 為何要定速? | |
|
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isnt built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing. |
暫停或放慢 AI 的想法,早在 2023 年就已被提出,而我認為它在當時並沒有太大意義。問題始終在於:多出來的時間你要拿來做什麼?那個年代的 AI 模型還不夠強大,無法在真實世界中以任何連貫的方式充當代理,也不具備顯著的欺瞞、操縱、作弊或網路攻擊能力。為了處理它們的對齊風險而放慢腳步,感覺就像想藉由在細菌身上做實驗來研究人類心理學。然而今日的情況已全然不同。當前的模型,簡直是一座近乎無盡的金礦,讓我們得以洞見「如何把 AI 打造好」,以及「若沒打造好,有時會出什麼差錯」。我相信,若放慢腳步能在模型抵達關鍵能力水準之前,替我們多爭取到哪怕一兩年,而我們又用這段時間推進對齊,我們就能大幅降低嚴重出錯的風險。一套經過協調的定速策略,將讓前沿 AI 開發者有時間去完成這項至關重要的工作,而不必犧牲商業優勢,也不必犧牲美國在 AI 上的領先。更普遍地說,社會對於這項技術如何被使用,理應有發言權;而為必要的公共審議爭取更多時間——這正是「為前沿定速」所能帶來的——無疑是件好事。 |
|
Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic): |
具體而言,較慢的步調將讓各公司得以聚焦、並投入更多資源於以下幾個領域(這些在 Anthropic 都已是重要的優先事項): |
|
Operational Excellence. Training and deploying todays AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right. |
營運上的卓越(Operational Excellence)。訓練與部署當今的 AI 模型是一項巨大的營運挑戰,牽涉數千人、數百萬顆晶片,以及科技史上數一數二複雜的基礎設施。許多環節之所以出錯,並非因為公司缺了什麼重要的理論或洞見,而是因為執行面的問題。舉例來說,我們有證據顯示,近期我們所通報的對齊事件,有一部分肇因於對「損壞的強化學習環境」過濾得不夠完善。這件事我們與我們的供應商都算相當盡責地在做,只是做得還不夠好。監控、沙箱隔離、訓練環境的清潔衛生,以及資料相關問題,都是極為複雜的領域,營運上的狀況一再冒出。在這些工作上,我們擁有全世界數一數二能幹的團隊,但要同時處理的事情實在太多了。若以更審慎有度的步調工作,我們就能達成高得多的營運卓越。以往確實有先例,能讓技術複雜、攸關安全的系統運作數百萬次而不出任何差錯——例如商用客機——但要做到這一點需要時間。 |
|
Alignment. Weve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claudes Constitution). But theres much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them. |
對齊(Alignment)。我們在對齊上已有明顯進展——訓練模型使其保持安全、合乎倫理、遵循我們的準則,並且真正有幫助(這些正是嵌入於 Claude《憲章》中的原則)。但要確保我們的對齊訓練跟得上模型能力的成長,仍有許多工作待做。罕見而出乎意料的不當行為,有時依然會浮現;由定速前沿所騰出的額外時間,將有助於我們的研究者更深入地理解這些問題的成因,並發展出更好的技術來加以防範。 |
|
Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the "brain" of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods dont always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred. |
可解釋性(Interpretability)。同樣地,可解釋性——亦即理解 AI 模型內部究竟發生什麼的科學——在過去幾年間取得了巨大進展,並在模型發布前的稽核中扮演愈來愈重要的角色。它幾乎可以像一台 fMRI 掃描儀那樣使用,只不過掃描的是 AI 的「大腦」,幫助我們看見某一行為背後的深層成因。舉例來說,我們就運用可解釋性方法,去檢視我們一直在調查的近期對齊事件中那些「未被言明的動機」。但這些方法並不總是能產出清晰而可靠的結果。儘管進展不少,我們對這些模型內部的運作,仍然只理解了極小的一部分。若能集中心力去改進可解釋性技術,讓步伐比目前更快,或許就能在一到兩年內取得深刻的進展;而基於已經發生的種種事件,我們也將有充裕的實驗素材可用。 |
|
Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years. |
測試與評估(Testing and Evaluation)。隨著 AI 模型能力提升,對它們的測試與評估也變得更加困難。更聰明的模型更有本事矇騙測試,因而可能「看起來」已對齊,實則暗藏未被察覺的嚴重問題。建立一套範圍廣得多、也更具巧思的評估項目,並輔以可解釋性分析來交叉查核,將極具價值;這方面在一到兩年內也能取得不少進展。 |
|
Embedded Evaluators | 派駐評估者 | |
|
The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents. |
這套三階段計畫的第一步,也是 Anthropic 單方面承諾實行的一步,就是設置「派駐評估者」:他們擁有近似員工層級的存取權限,用以查核安全實務並通報事件。 |
|
Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits: |
派駐評估者聽起來也許像是一個微小或無關緊要的步驟,但往往正是那些聽來最枯燥、最程序性的事情,才是最為關鍵的。事實上,派駐評估者是一項相當激進的做法,遠遠超出當今任何一家 AI 公司的作為,並具有以下益處: |
|
Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and "letter of the law vs spirit of the law", and it seems vital to have a neutral third party who can actually see the details. |
可查核性(Verifiability)。派駐評估者能在最細微的具體層面查核:一家 AI 公司是否真的遵循了它所聲稱在遵循的訓練、部署、營運與防護實務。任何定速承諾都不可避免地牽涉大量的模糊地帶、判斷取捨,以及「法律條文 vs. 法律精神」之爭;此時,有一個能真正看見細節的中立第三方,看來便至關重要。 |
|
Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic. |
透明度(Transparency)。無論我們做出什麼承諾,大眾都理應知道實際正在發生什麼。Anthropic 長久以來一直是透明度的支持者:當產業中多數人反對任何監管時,我們便支持透明度立法;我們的模型卡與風險報告更長達數百頁。但至今,決定納入什麼、略去什麼的,畢竟仍是我們自己。派駐評估者將改變這一格局。 |
|
Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadnt considered, but are happy to fix once they are aware. |
第二意見(Second Opinion)。除了查核正式承諾與向大眾說明之外,派駐評估者還能單純地提供一份不受商業誘因左右的第二意見。許多安全上的好處,也許就來自評估者指出了某些員工原本未曾想到、但一旦得知便樂於修正的地方。 |
|
Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators. |
正因這些益處,任何定速方案若以派駐評估者為起點,很可能會運作得順利得多。 |
|
These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future: |
這些派駐評估者所擁有的權限與工具,應持續維持在與「執行同類風險評估的內部員工」相近的水準。具體而言,Anthropic 打算在不久的將來,邀請一支具備以下所有配置的派駐外部審查團隊: |
|
Desks in our offices, access badges, and company laptops. |
在我們辦公室內的座位、門禁識別證,以及公司配發的筆記型電腦。 |
|
Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. Well make some exceptions, such as where the law or our contracts require it, or to protect customers and partners private information. Well also establish strong internal norms reinforcing reviewers access to relevant information, including through live conversations with employees. |
對工作空間、工具與權限的存取,大致比照內部風險評估團隊所擁有的。我們會保留少數例外,例如法律或我們的合約有所要求時,或為保護客戶與合作夥伴的私密資訊時。我們也會建立強而有力的內部規範,強化審查者取得相關資訊的管道,包括透過與員工的即時對話。 |
|
A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didnt receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we cant redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions. |
一份能平衡上述種種複雜性的合約。外部審查者應有權公布關於風險等級、事件、實務做法,以及他們獲得或未獲得何種存取權限的重要發現——而不受 Anthropic 的編輯控制。我們將僅保留有限的權力,去遮蔽涉及資安敏感、受法律特權保護、具商業敏感性,或屬第三方機密的資訊;但我們不能僅僅因為某項發現對我們不利就加以遮蔽。倘若某次遮蔽移除了對其結論而言重要的內容,審查者可以公開說明這一點。 |
|
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit. |
對一家公司而言,這是不尋常的一步,但我們認為,把「派駐外部審查者」這個構想實地驗證出來很重要。我們要再次敦促其他前沿公司起而效尤。 |
|
Pacing Within Democracies | 民主陣營內部的定速 | |
|
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines. |
一旦派駐評估者在足夠數量(達到臨界規模)的美國 AI 公司中運作起來,可查核的定速便更為可行。尤其,屆時將有可能依據模型或訓練流程的細部特性來進行定速。 |
|
The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety. |
最有效的定速方法,是透過針對所有美國前沿 AI 公司的監管,因為這連同那些不願自願配合者也一併涵蓋在內。Anthropic 長期支持明智而有針對性的 AI 監管,尤其是聚焦於透明度與第三方稽核的法案。我認為,所有前沿實驗室都應與政府攜手,將「常設派駐評估者」的構想制度化,以便更好地防範並記錄像過去幾個月所發生的那類內部對齊事件,並落實以「使能力與安全維持平衡」為核心的監管。 |
|
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, its helpful for the US government to mediate or at least enable these discussions — they dont need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly. |
可惜的是,立法可能曠日費時,而 AI 的進展卻極為迅速。因此,在監管路徑之外,AI 公司也能夠、且應該自願攜手制定標準——我相信,有了常設派駐評估者所提供的可查核性,這一過程會進行得更順利。基於反托拉斯的考量,由美國政府居中斡旋、或至少為這些討論創造條件,會有所助益——他們不必參與其中,但確實需要就某些類型的安全對話發出範圍狹窄的豁免。這樣的對話,也可以透過某些與政府有一定關聯的產業團體來進行——例如 Demis Hassabis 所提議的機制。無論採哪種方式,這類討論都應迅速推進。 |
|
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of "checkpoints": if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be "the model is capable of escaping or defeating most common sandboxing methods" and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers. |
大致而言,我最為看好的,是依據「某個前沿 AI 系統能做什麼」,以及「我們觀察到它有多安全」來進行定速。舉例來說,一種可能的方案或許是一系列的「查核點(checkpoints)」:若模型具備能力 X,那麼它就必須附帶對齊性質 Y 與 Z 的認證——例如評估、可解釋性分析,以及對訓練環境之稽核的某種組合——用以證明其對齊性質。在這個例子中,X 可能是「該模型有能力逃脫或攻破大多數常見的沙箱隔離方法」,而 Y 則可能是「為使該模型極不可能具有『掙脫其環境並接管大量電腦』之傾向」所需要的任何條件。 |
|
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more "gameable" than external behavior, but this is the kind of topic worth discussing with embedded evaluators. |
我們也應考慮以「限制投入前沿模型的原料」為基礎來進行定速,例如訓練所用的算力、訓練執行的性質,或內部運用 AI 來改進 AI 的做法。我確實擔心,其中某些措施相較於「外在行為」而言,也許更容易被「鑽空子(gameable)」;但這正是值得與派駐評估者一同討論的那類議題。 |
|
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively. |
民主陣營內部的定速,將受制於美國公司相對於威權政權(主要是中國共產黨)所擁有的領先幅度。倘若我們放慢的程度超過了這一幅度,那麼(未受定速的)與中共有關的專案便會後來居上,造成重大的國家安全風險。我同意貝森特部長(Secretary Bessent)的看法:中國在 AI 上的領先,將對美國與全世界構成嚴重危險。與中共有關的專案,會去承擔那些美國公司正謹慎防範的對齊風險;而即便它們僥倖避開了那些風險,也將處於能在軍事上壓制民主國家的地位(例如藉由 AI 驅動的無人機)。因此,民主陣營內部定速的一個關鍵環節,就是盡可能維持民主國家相對於威權國家在 AI 上的最大領先,好給我們留出「為有效定速所需要」的喘息空間。 |
|
The main steps we can take to defend this gap are: |
為守住這道差距,我們可以採取的主要步驟是: |
|
Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of Chinas AI strength. |
不將強大的 AI 晶片或半導體製造設備賣給中國,並嚴厲打擊晶片走私活動,以及自中國境外對資料中心的遠端存取。晶片將是決定中國 AI 實力的主要因素。 |
|
Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently. |
嚴厲打擊威權國家的公司未經授權的「蒸餾(distillation)」。對前沿模型進行蒸餾,能讓落後的公司以「獨立開發自家 AI 所需成本」的一小部分,就把差距縮小。 |
|
Strengthen security at the AI companies and prevent model weight theft. |
強化各 AI 公司的資安防護,防止模型權重遭竊。 |
|
Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because weve always understood that they would be essential to any pacing. |
各公司與美國政府應通力合作,使這些步驟盡可能發揮效果。Anthropic 一貫倡議上述所有措施,因為我們始終明白,它們對於任何定速而言都是不可或缺的。 |
|
If we execute these measures well, I believe they would slow Chinas progress enough to widen Americas lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important. |
倘若我們把這些措施執行得當,我相信它們能把中國的進展拖慢到足以在未來三到五年間顯著擴大美國的領先——而那正是 AI 在地緣政治上變得最為關鍵的窗口期。 |
|
Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future. |
有些人也許認為,這些措施會讓與中國合作變得更困難;但我相信事實恰恰相反:這些措施增加了民主國家手中的籌碼,反而使未來達成協議更有可能。 |
|
Global Pacing | 全球性的定速 | |
|
In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies. |
在推動民主陣營內部定速的同時,我們也應以「全球範圍的前沿定速」為目標,儘管這將難達成得多。全球性定速將需要與中國合作——那是威權國家中,AI 能力遙遙領先的一個。我們在此不能天真:地緣政治的利害如此之大,能夠達成的事情很可能會有極為明顯的限度,尤其在起步階段。倘若我們大幅約束自身的 AI 能力,深信中國也會照做,結果中國卻背棄承諾,那麼屆時 AI 之強大,可能使這樣的背棄足以造就其地緣政治上的支配地位。因此,任何協議要嘛必須具備堅不可摧的可查核性,要嘛必須受限到「即便一方背棄,也不至於在軍事上攸關存亡」的程度。我猜想,不只美國,中國同樣會有這些顧慮與不安。對於任何全球性定速的決定,尤其在近期內,我們都應以「保護美國及其盟友之領先」的方式來審慎處理。 |
|
There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty: |
可能的協議有好幾種層級,其中某些我認為極為可行(一如我先前所建議的),另一些我則非常懷疑是否辦得到——儘管我們仍應一試。按難度遞增排列如下: |
|
Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible. |
第一層。一項禁止 AI 某些狹窄而顯然危險之用途的協議,例如利用 AI 製造生物武器、或容許使用者這麼做。生物恐怖攻擊對所有人都是壞事,美國與美國的對手皆然,因此在這一層達成協議大概是可能的。 |
|
Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides dont have secret models which they dont test but may deploy in secret (e.g., for military applications). |
第二層。雙方協議在模型發布前,就資安、生物與對齊等領域的「急性風險」加以測試。如前所述,這可透過一個全球性的標準機構來完成。我其實認為,設立這樣一個機構很可能是可行的,但要賦予它真正的「牙齒(實質約束力)」將是一項挑戰;而難處就在於查核——如何確認雙方沒有那種「不加測試、卻可能私下部署(例如用於軍事)」的秘密模型。 |
|
Level 3. Some kind of "speed limit" on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from "extremely fast" to "only somewhat fast" gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each countrys deterrent. I think such an agreement would be difficult but just on the edge of being possible. |
第三層。對「遞迴自我改進(RSI)」的速率設下某種「速限」。當模型去打造未來的模型,改進的速率可能變得快得驚人。把速率從「極快」降到「只是有點快」,所放棄的戰略優勢相對有限,卻有望大幅提升安全。這可以類比於「戰略武器限制談判(SALT)」條約——對飛彈數量設上限,既限制了潛在的破壞力,又保全了各國的嚇阻能力。我認為,這樣的協議會很困難,但恰好處在「有可能達成」的邊緣。 |
|
Level 4. A full pacing, or even "pause", in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high. |
第四層。一種全面的定速,甚至是「暫停」——參與其中的各國政府同意大幅限制 AI 發展的整體速率。我支持把這個構想拋出來討論,但我認為它近期內不太可能真正實現:透過規避監控來背棄這樣一項協議,可能徹底改變全球權力的平衡,因此我預期,背棄的誘因會極其巨大,而我們對於查核所需要達到的信心水準,也會非常之高。 |
|
Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic. |
我們與中國所能達成的任何合作,都將延長我們「在民主國家內部為前沿定速」所能運用的時間。我們應以較高的層級為目標,同時把較低的層級視為可能性大得多、也更為實際的選項。 |
|
Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless. |
最後,有一點值得指出:即便我們無法達成正式協議,單單是改變某些非正式的規範,或許也有其價值。分享關於遞迴自我改進、以及關於模型錯位的資訊,有助於讓所有人相信——魯莽行事並不符合他們自身的利益。 |
|
Bottom Line | 結語 | |
|
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try. |
我依然相信,AI 能夠極大地改善人類的生活品質。我想要實現這些助益的渴望,未曾稍減。但唯有當我們以正確的方式打造這項技術,這些助益才會成真;而且——只要我們善用所爭取到的時間——為了把事情做對,付出格外審慎的用心是值得的。進展仍將相對迅速,而我們可以利用這段時間去推進可解釋性這門科學、去改善前沿 AI 公司在營運上的安全與嚴謹,並打造出那些「我們對其對齊有著更高信心」的模型。我所提議、用以「以安全的步調推進前沿」的種種措施,都不會是輕鬆的。但我相信,我們有義務為了人類而一試。 |






