Skip to main content
Aggregate AI 摘要 arXiv cs.AI 人工智能 7 Sep 2026 - 15:00

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

RSS 官方收录 · 可信分层展示

关键摘要

C-DPPO新算法提升编码模型性能3.0分,保障训练与部署token一致性

  • 提出 fidelity-aware 训练框架,确保原始prompt端到端采样
  • 首创C-DPPO算法,提供TV距离双向认证与错误鲁棒策略掩码
  • 在Baize5B/10B上实测TMax-100任务,稳定+3.0分超越标准DPPO

AI 摘要 · 来源可核验

正文提要

arXiv:2609.04678v1 Announce Type: new Abstract: Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates spurious model calls via a negotiated training protocol, and restricts loss computation to verifiable token spans with closed-failure guarantees. We further propose Certified Divergence Proximal Policy Optimization (C-DPPO), which establishes tight two-sided TV certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking atop standard DPPO. Evaluated on matched Baize5B and Baize10B models with identical training and test protocols on TMax-100, C-DPPO yields a consistent +3.0-point performance gain over standard DPPO across model scales. Certificate audits validate the reliability and full operational coverage of our certified training pipeline.

来源:https://arxiv.org/abs/2609.04678

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表