Diluner/gpt54-mini-standalone-qwen3-4b-rose-babyai-20260920
Diluner/gpt54-mini-standalone-qwen3-4b-rose-babyai-20260920 is a 4 billion parameter Qwen3-based causal language model, trained independently with the ROSE method using a gpt-5.4-mini teacher. This model is specifically fine-tuned for the BabyAI environment, achieving an 87.5000% average success rate across four attempts per official test task. It is optimized for task completion within the BabyAI domain, demonstrating proficiency in navigating and interacting with simulated environments.
Loading preview...
Model Overview
Diluner/gpt54-mini-standalone-qwen3-4b-rose-babyai-20260920 is a 4 billion parameter model built on the Qwen3 architecture. It was trained using the ROSE (Reinforcement Learning from Optimal Self-Correction) method, guided by a gpt-5.4-mini teacher model. This particular checkpoint represents the final stage of a dedicated training process for the BabyAI environment, completed over five epochs.
Key Capabilities
- BabyAI Task Proficiency: Achieves an 87.5000% average success rate on BabyAI tasks, evaluated across four attempts per official test task.
- Standalone Training: The model was initialized and trained independently, rather than as a sequential checkpoint of a base model, indicating a focused optimization for its target domain.
- Robust Evaluation: Evaluation metrics include mean success across four attempts (
avg@4) and verified task/sample coverage, with zero episode errors recorded during testing.
Good For
- Research in Reinforcement Learning: Ideal for researchers exploring the effectiveness of the ROSE training method and teacher-student learning paradigms.
- BabyAI Environment Development: Suitable for developers and researchers working on agents for the BabyAI platform, offering a strong baseline for task completion.
- Understanding Specialized Fine-tuning: Provides an example of a model highly specialized for a particular simulated environment, showcasing the impact of targeted training.