Skip to content

What’s the difference between this and the slime framework? Other than support for a LoRA feature. But it seems that Miles supports LoRA RL. #1

Description

@yefei12
No description provided.

Activity

  1. changed the title [-]What’s the difference between this and the slime framework? Other than support for a LoRA feature.[/-] [+]What’s the difference between this and the slime framework? Other than support for a LoRA feature. But it seems that Miles supports LoRA RL.[/+] on May 28, 2026
  2. zqiu24 commented on May 28, 2026

    @zqiu24
    Contributor

    Thanks for the great question!

    Our motivation for starting Orbit was simple: we wanted single-node RL post-train models such as Kimi K2.6 and DeepSeek V4 with Orthogonal Finetuning / Quantized Orthogonal Finetuning (OFT / QOFT) [1,2,3]. At the time, we did not find an existing framework that supported this workflow directly. We were inspired by many excellent open-source RL frameworks, including miles, verl, and slime, and we built on ideas from the broader community.

    When post-training trillion-parameter models on a single node, we have to rely on extremely low-bit quantized base models (like we did with Kimi-K2.6 and DeepSeek V4-Pro). To enable single-node RL training of models on such a scale, we directly finetune the quantized weights without the overhead of storing a dequantized base. This is an area where we have invested significant engineering effort, for example, in ensuring that the log-prob diff remains stable even when finetuning trillion-scale quantized base models.

    Beyond hardware constraints, training stability remains the primary bottleneck at this scale. Standard QLoRA frequently struggles to maintain stability during rigorous RL loops, which is exactly why a robust, alternative approach is essential. Our evaluations have shown that OFT consistently demonstrates much better training stability and general capability preservation than LoRA in these low-precision RL settings. This is why it has been chosen as the default PEFT optimizer.
    Orbit’s architecture is built from the ground up to deliver the same system-level scalability and memory footprint advantages as QLoRA, while leveraging QOFT's mathematical stability, eliminating the precision mismatch between the training and serving base, and ensuring even high-throughput training with optimized OFT kernels.

    Thanks again for checking out the project! Let us know if you have any other questions.

    [1] Controlling Text-to-Image Diffusion by Orthogonal Finetuning (Qiu, Liu, et al, NeurIPS 2023)
    [2] Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization (Liu, Qiu, et al, ICLR 2024)
    [3] Orthogonal Finetuning Made Scalable (Qiu, Liu, et al, EMNLP 2025)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions