Aeronix-zzz/VERSE-rrs-7B
Aeronix-zzz/VERSE-rrs-7B is a 7 billion parameter language model, representing Stage 2 of the VERSE project. It is a continue-SFT model based on filtered self-correction trajectories, serving as the initial policy for the self-evolve GRPO stage. This model is designed for tasks requiring reflective reasoning and self-correction, building upon its cold-start SFT predecessor. It is particularly suited for applications involving VQA, maze solving, and scene understanding where iterative refinement is beneficial.
Loading preview...
VERSE-rrs-7B: Reflective Rejection-Sampling SFT Model
VERSE-rrs-7B is the second stage in the VERSE project by Aeronix-zzz, a 7 billion parameter model focused on enhancing reasoning through self-correction. It is a continue-SFT (Supervised Fine-Tuning) model, specifically trained on filtered self-correction trajectories. This training methodology involves a process of 'think→draw→observe→revise', allowing the model to iteratively refine its outputs.
This model serves as the crucial initial policy for the subsequent self-evolve GRPO (Generative Reinforcement Learning with Policy Optimization) stage, indicating its role in developing more advanced, self-improving AI systems.
Key Capabilities
- Reflective Reasoning: Trained on self-correction trajectories, enabling iterative refinement of responses.
- Foundation for Advanced Stages: Acts as the base policy for the self-evolve GRPO stage of the VERSE project.
- Specialized Training Data: Utilizes diverse datasets including
vqa,maze,GPT4Scene, andSR_91kreflective data for its rejection-sampling SFT.
Good for
- Research in Self-Correction: Ideal for developers and researchers exploring models capable of iterative self-improvement.
- Complex Reasoning Tasks: Suitable for applications requiring reflective thinking, such as visual question answering (VQA) and maze navigation.
- Scene Understanding: Can be applied to tasks involving detailed analysis and understanding of scenes, leveraging its specialized training data.