遵守莫训
07-02 · 中船重工
系统架构
ʻO ka Helu
Setup: Let pi_theta be a Turing-complete policy over MDP mathcal{M}, and H(theta_1,theta_2) a human preference oracle. Define the RLHF objective:J(theta) = mathbb{E}{tausimpitheta}left[sum gamma^t R_tright] - beta D_{KL}(pi_theta | pi_{ref})Theorem (Alignment Incompleteness): There is no algorithm that decides, for arbitrary theta, whether nabla J(theta)=0 implies pi_theta is provably safe w.r.t. H, unless ZFC is inconsistent.Proof Sketch: Turing-completeness makes halting Pi_2^0-complete (Rice). Safety requires universal quantification over infinite trajectories (Pi_1^1). By Gödel, the statement “exists theta^: nabla J(theta^)=0 land text{safe}(theta^*)” is independent of ZFC+RLHF-convergence. Thus, critical points ≠ provable alignment. ∎Corollary: PPO converges to a fixed point. Whether it’s the safe one is formally unprovable within your training axioms. Restrict expressivity or accept undecidability. No third option, darling.(Word count: 148)
发布于 河北
分享
评论
赞
未登录
友善发言
评论
加载中
下载脉脉APP,成就职业梦想
违法不良信息&未成年人有害信息举报电话/客服电话:400 065 0808
违法不良信息&未成年人有害信息举报邮箱/客服邮箱:maimai@taou.com
清朗系列专项行动相关违规信息举报电话:400 065 0808,举报邮箱:maimai@taou.com
个人/企业等被诽谤侮辱、人身权或知识产权等被侵犯、网络谣言的举报地址:maimai.cn/tousu | 涉企虚假不实信息举报投诉专区
京ICP备12005786号-1copyright©maimai.cn