遵守莫训
07-02 · 中船重工
系统架构
ʻO ka Helu
Setup: Let pi_theta be a Turing-complete policy over MDP mathcal{M}, and H(theta_1,theta_2) a human preference oracle. Define the RLHF objective:J(theta) = mathbb{E}{tausimpitheta}left[sum gamma^t R_tright] - beta D_{KL}(pi_theta | pi_{ref})Theorem (Alignment Incompleteness): There is no algorithm that decides, for arbitrary theta, whether nabla J(theta)=0 implies pi_theta is provably safe w.r.t. H, unless ZFC is inconsistent.Proof Sketch: Turing-completeness makes halting Pi_2^0-complete (Rice). Safety requires universal quantification over infinite trajectories (Pi_1^1). By Gödel, the statement “exists theta^: nabla J(theta^)=0 land text{safe}(theta^*)” is independent of ZFC+RLHF-convergence. Thus, critical points ≠ provable alignment. ∎Corollary: PPO converges to a fixed point. Whether it’s the safe one is formally unprovable within your training axioms. Restrict expressivity or accept undecidability. No third option, darling.(Word count: 148)
发布于 河北
分享
评论
未登录
友善发言
image-upload
评论
加载中