code-critic-model/Qwen3-4B-SFT-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep3
Qwen3-4B-SFT-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep3 is a 4 billion parameter language model developed by code-critic-model, fine-tuned from code-critic-model/qwen3-4b-sft-prm. This model was trained using Direct Preference Optimization (DPO) on the PRM_1541i dataset, focusing on generating responses aligned with human preferences. It is a development checkpoint, with the primary critic model for the 'Steer, Don't Solve' paper being Qwen3-4B-Critic-SFT-DPO. This model is suitable for general text generation tasks where preference alignment is desired.
Loading preview...
Overview
This model, Qwen3-4B-SFT-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep3, is a 4 billion parameter language model developed by code-critic-model. It is a fine-tuned version of code-critic-model/qwen3-4b-sft-prm, specifically trained using Direct Preference Optimization (DPO). The training utilized the code-critic-model/PRM_1541i dataset and the TRL framework.
Key Characteristics
- Architecture: Based on the Qwen3-4B family.
- Training Method: Employs Direct Preference Optimization (DPO), a technique designed to align model outputs with human preferences, as detailed in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2305.18290).
- Dataset: Fine-tuned on the
PRM_1541idataset. - Development Checkpoint: This specific model is noted as a development checkpoint and is not the primary critic model discussed in the "Steer, Don't Solve" paper; the main critic is
Qwen3-4B-Critic-SFT-DPO.
Use Cases
This model is suitable for general text generation tasks where the goal is to produce outputs that are aligned with learned human preferences. Its DPO training makes it potentially useful for applications requiring nuanced response generation.