Metaphor
Search
搜索
暗色模式
亮色模式
探索
标签: grpo
此标签下有10条笔记。
2026年6月20日
强化学习后训练专题索引
reinforcement-learning
llm-training
grpo
index
2026年6月20日
大型推理模型的自我改进技术 (HSIR)
llm
reasoning
self-improvement
grpo
2026年6月20日
高级GRPO变体综述:Latent-GRPO、SPPO、BPPO与LamPO
reinforcement-learning
grpo
reasoning
efficiency
2026年6月20日
f-GRPO:基于散度的强化学习统一框架
reinforcement-learning
llm-alignment
grpo
dpo
divergence
2026年6月20日
GRPO-VPS:可验证过程监督增强的组相对策略优化
reinforcement-learning
grpo
process-supervision
reasoning
2026年6月20日
PPO、GRPO与DAPO算法对比分析
reinforcement-learning
llm-training
grpo
dapo
ppo
2026年5月17日
GRPO理论基础与LLM对齐
reinforcement-learning
grpo
llm-alignment
policy-gradient
2026年5月08日
Elastic Reasoning:可扩展的思维链推理框架
elastic-reasoning
chain-of-thought
reasoning-models
inference-time-scaling
grpo
test-time-compute
2026年5月03日
dUltra:强化学习加速扩散语言模型
diffusion
reinforcement-learning
language-model
grpo
2026年5月01日
推理模型架构
reasoning-models
o1
r1
openai
deepseek
grpo
reinforcement-learning