Diluner/gpt54-mini-sequential-qwen3-4b-rose-s1-babyai-20260920

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026Architecture:Transformer Featherless Exclusive Cold

Diluner/gpt54-mini-sequential-qwen3-4b-rose-s1-babyai-20260920 is a 4 billion parameter Qwen3-based causal language model, fine-tuned using the ROSE method with a gpt-5.4-mini teacher. This specific checkpoint represents the completion of the first sequential stage (BabyAI) of a multi-stage training process. It is designed for sequential learning environments, demonstrating a specialized training methodology rather than general-purpose instruction following.

Loading preview...

Model Overview

This model, gpt54-mini-sequential-qwen3-4b-rose-s1-babyai-20260920, is a 4 billion parameter variant of the Qwen3 architecture. It has undergone a specialized training regimen using the ROSE method, with gpt-5.4-mini serving as the teacher model. This particular release is an intermediate checkpoint, marking the successful completion of the first sequential stage (BabyAI) in a planned multi-stage training pipeline that includes TextCraft and SearchQA.

Key Training Characteristics

  • Teacher-Student Learning: Utilizes gpt-5.4-mini as a teacher to guide the Qwen3-4B student model's learning process.
  • Sequential Training: Represents the final checkpoint of the BabyAI stage, which involved five epochs and 125 optimizer updates for that specific environment.
  • Intermediate Checkpoint: This is not the final, fully trained model across all stages. Evaluations for later sequential checkpoints or standalone models should not be attributed to this specific checkpoint.
  • Provenance: Part of a distinct sequential training chain, not a historical standalone run, with each stage building upon the preceding one.

Usage Notes

This model is provided as a trained checkpoint from a specific experimental setup. It is important to note that:

  • No completed evaluation scores are attached to this intermediate checkpoint.
  • It serves as evidence of a particular training methodology rather than a general-purpose, instruction-tuned model.
  • The repository includes the model configuration, tokenizer, and all weight shards, but excludes optimizer state, raw logs, and teacher trajectories.

Developers can load the model using the Hugging Face transformers library:

from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "Diluner/gpt54-mini-sequential-qwen3-4b-rose-s1-babyai-20260920"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, torch_dtype="auto")