code-critic-model/Qwen3-4B-SFT-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026Architecture:Transformer Featherless Exclusive Cold

This 4 billion parameter Qwen3-based model, developed by code-critic-model, is a fine-tuned language model with a 32768 token context length. It was trained using Direct Preference Optimization (DPO) on the PRM_1541i dataset, making it suitable for tasks requiring preference-based learning and refined text generation.

Loading preview...

Model Overview

This model, Qwen3-4B-SFT-DPO-beta0.1-sft0.25-lr1e-6-bs32-ep1, is a 4 billion parameter language model based on the Qwen3 architecture. It was developed by code-critic-model and fine-tuned from code-critic-model/qwen3-4b-sft-prm.

Key Capabilities

  • Preference-based Learning: The model was trained using Direct Preference Optimization (DPO), a method that leverages human preferences to refine model outputs, as detailed in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model".
  • Dataset Specificity: Fine-tuned on the code-critic-model/PRM_1541i dataset, suggesting potential strengths in areas related to this dataset's content.
  • Framework: Training was conducted using the TRL (Transformers Reinforcement Learning) library, indicating a focus on reinforcement learning from human feedback (RLHF) techniques.

Good For

  • Research in DPO: Ideal for researchers exploring the effects and applications of Direct Preference Optimization on Qwen3-based models.
  • Preference-aligned Text Generation: Suitable for tasks where aligning model outputs with specific preferences is crucial, potentially leading to more nuanced or desired responses.
  • Experimental Use: As a development checkpoint, it serves as a valuable resource for understanding the evolution of code-critic-model's DPO-trained models, particularly in the context of the PRM_1541i dataset.