You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
What’s the difference between this and the slime framework? Other than support for a LoRA feature. But it seems that Miles supports LoRA RL. #1
changed the title [-]What’s the difference between this and the slime framework? Other than support for a LoRA feature.[/-][+]What’s the difference between this and the slime framework? Other than support for a LoRA feature. But it seems that Miles supports LoRA RL.[/+]on May 28, 2026
Our motivation for starting Orbit was simple: we wanted single-node RL post-train models such as Kimi K2.6 and DeepSeek V4 with Orthogonal Finetuning / Quantized Orthogonal Finetuning (OFT / QOFT) [1,2,3]. At the time, we did not find an existing framework that supported this workflow directly. We were inspired by many excellent open-source RL frameworks, including miles, verl, and slime, and we built on ideas from the broader community.
When post-training trillion-parameter models on a single node, we have to rely on extremely low-bit quantized base models (like we did with Kimi-K2.6 and DeepSeek V4-Pro). To enable single-node RL training of models on such a scale, we directly finetune the quantized weights without the overhead of storing a dequantized base. This is an area where we have invested significant engineering effort, for example, in ensuring that the log-prob diff remains stable even when finetuning trillion-scale quantized base models.
Beyond hardware constraints, training stability remains the primary bottleneck at this scale. Standard QLoRA frequently struggles to maintain stability during rigorous RL loops, which is exactly why a robust, alternative approach is essential. Our evaluations have shown that OFT consistently demonstrates much better training stability and general capability preservation than LoRA in these low-precision RL settings. This is why it has been chosen as the default PEFT optimizer.
Orbit’s architecture is built from the ground up to deliver the same system-level scalability and memory footprint advantages as QLoRA, while leveraging QOFT's mathematical stability, eliminating the precision mismatch between the training and serving base, and ensuring even high-throughput training with optimized OFT kernels.
Thanks again for checking out the project! Let us know if you have any other questions.