shubhamrgandhi/qwen3-4b-instruct-2507-full-sft-prm-r2egym-swebench-instructions-k5-qwen-only
This model is a fine-tuned Qwen3-4B-Instruct-2507, a 4 billion parameter instruction-tuned causal language model developed by Qwen, with a context length of 32768 tokens. It has been specifically fine-tuned on the prm_sft_train dataset. This model is designed for general instruction-following tasks, leveraging its base Qwen3 architecture and specialized fine-tuning.
Loading preview...
Model Overview
This model, shubhamrgandhi/qwen3-4b-instruct-2507-full-sft-prm-r2egym-swebench-instructions-k5-qwen-only, is a specialized fine-tuned version of the Qwen3-4B-Instruct-2507 base model, developed by Qwen. It features 4 billion parameters and supports a substantial context length of 32768 tokens, making it suitable for processing longer inputs and generating comprehensive responses.
Key Characteristics
- Base Model: Qwen3-4B-Instruct-2507, a robust instruction-tuned causal language model.
- Fine-tuning: Specifically trained on the
prm_sft_traindataset, indicating a focus on particular instruction-following or problem-solving tasks, though specific details of the dataset's content are not provided in the README. - Parameter Count: 4 billion parameters, offering a balance between performance and computational efficiency.
- Context Length: 32768 tokens, enabling the model to handle extensive conversational histories or detailed documents.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 5e-06
- Optimizer: ADAMW_TORCH_FUSED
- Epochs: 3.0
- Batch Size: A total training batch size of 8 across 8 GPUs.
Intended Use Cases
Given its instruction-tuned nature and fine-tuning, this model is generally suitable for:
- Instruction-following tasks.
- Generating text based on specific prompts.
- Applications requiring a model with a large context window for detailed interactions.