Diluner/gpt54-mini-sequential-qwen3-4b-rose-s2-textcraft-20260920

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026Architecture:Transformer Featherless Exclusive Cold

Diluner/gpt54-mini-sequential-qwen3-4b-rose-s2-textcraft-20260920 is a 4 billion parameter Qwen3 causal language model, fine-tuned using the ROSE method with a gpt-5.4-mini teacher. This model represents the second stage (TextCraft) of a sequential training process, following an initial BabyAI stage. It is optimized for text generation tasks, having completed five epochs in the TextCraft environment, and is an intermediate checkpoint in a multi-stage training chain.

Loading preview...

Model Overview

This model, gpt54-mini-sequential-qwen3-4b-rose-s2-textcraft-20260920, is a 4 billion parameter Qwen3-based causal language model developed by Diluner. It has been trained using the ROSE method with gpt-5.4-mini as the teacher model. This particular release is the final checkpoint of the TextCraft stage, having completed five epochs and 55 optimizer updates within that stage.

Sequential Training Process

This model is part of a sequential training chain that progresses through different environments: BabyAI → TextCraft → SearchQA. This checkpoint specifically represents the completion of the TextCraft stage, following an initial BabyAI stage. Each environment in this sequence receives five epochs of training, with the student model and training method carried through the entire chain.

Key Characteristics

  • Architecture: Qwen3-4B, a 4 billion parameter causal language model.
  • Training Method: ROSE (Reinforced Optimization for Sequential Environments) with gpt-5.4-mini as the teacher.
  • Context Length: Supports a context length of 32768 tokens.
  • Training Stage: Represents the completed TextCraft stage (stage 2) of a three-stage sequential training process.
  • Intermediate Checkpoint: This is an intermediate checkpoint; final evaluation scores for the full three-environment sequence are not attached to this specific release.

Intended Use

This model is suitable for developers interested in exploring models trained with sequential, multi-stage methods, particularly those focusing on text generation tasks as indicated by its TextCraft stage completion. As an intermediate checkpoint, it provides a snapshot of the model's capabilities after specific training phases. Users should note that comprehensive evaluations are intended for the final, fully trained stage-3 model.