Hugging Face researchers describe a method for running asynchronous GRPO (Group Relative Policy Optimization) with LoRA (Low-Rank Adaptation) across multiple HF Jobs. The key innovation is decoupling the update cycle from the usual all-to-all communication required by NCCL, replacing it with a lightweight coordination mechanism built around a bucket and a proxy. This allows each worker to proceed independently and periodically sync its progress.

The practical payoff is that reinforcement learning fine-tuning of large models becomes easier to scale across heterogeneous clusters. Because LoRA only updates a small set of adapters, the per-step memory and compute cost stay low, and the asynchronous design means a slow node does not stall the entire job. The authors report this setup works across separate HF Jobs without needing to configure NCCL, simplifying deployment.

While the blog does not provide benchmark numbers or comparisons to synchronous methods, it outlines a concrete recipe that other teams can adopt. The approach is particularly relevant for teams already using Hugging Face's training infrastructure and looking to experiment with policy optimization without investing in specialized networking. The tradeoff is that asynchronous updates can introduce staleness, which may affect convergence speed or stability in some settings.