shubhamrgandhi/qwen3-4b-instruct-2507-full-sft-prm-r2egym-swebench-instructions-k5-cwm-plus-qwen

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

This model is a 4 billion parameter instruction-tuned Qwen3 model, fine-tuned from Qwen/Qwen3-4B-Instruct-2507. It is optimized for specific instruction-following tasks, leveraging a 32K context length. The model is designed for applications requiring robust instruction adherence and general language understanding.

Loading preview...

Model Overview

This model, shubhamrgandhi/qwen3-4b-instruct-2507-full-sft-prm-r2egym-swebench-instructions-k5-cwm-plus-qwen, is a specialized instruction-tuned variant of the Qwen3-4B-Instruct-2507 base model developed by Qwen. It has been fine-tuned on the prm_sft_train dataset, enhancing its ability to follow instructions effectively. With 4 billion parameters and a context length of 32,768 tokens, it is suitable for a range of natural language processing tasks.

Key Training Details

The fine-tuning process involved specific hyperparameters to optimize performance:

  • Learning Rate: 5e-06
  • Optimizer: ADAMW_TORCH_FUSED
  • Epochs: 3.0
  • Distributed Training: Utilized 8 devices for multi-GPU training.

This configuration aims to improve the model's instruction-following capabilities and overall responsiveness to user prompts. The model leverages recent versions of key frameworks, including Transformers 5.2.0 and Pytorch 2.9.1+cu128.