cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed100-25grp
The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed100-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in the DeepSeekMath paper, to enhance its capabilities. It is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, leveraging its 32K context length. This model is suitable for applications demanding robust analytical and problem-solving skills.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed100-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL library.
Key Capabilities
- Enhanced Reasoning: This model incorporates the GRPO (Grouped Reinforcement Learning with Policy Optimization) training method, which was originally introduced in the DeepSeekMath paper. This method is designed to push the limits of mathematical reasoning in language models.
- Instruction Following: As a fine-tuned instruction model, it is capable of understanding and executing a wide range of user prompts.
- Extended Context: Features a substantial context length of 32,768 tokens, allowing it to process and generate longer, more complex sequences of text.
Training Details
The model's training leveraged specific versions of popular frameworks:
- TRL: 0.19.1
- Transformers: 4.57.6
- Pytorch: 2.5.1
- Datasets: 4.8.4
- Tokenizers: 0.22.2
Good For
- Applications requiring strong mathematical or logical reasoning.
- Tasks benefiting from advanced instruction following in a compact model size.
- Scenarios where processing longer inputs or generating detailed responses is crucial.