code-critic-model/Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-120

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026Architecture:Transformer Featherless Exclusive Cold

Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-120 is a 4 billion parameter language model developed by code-critic-model, serving as a development checkpoint from a DPO training run. This model is a later iteration of the Qwen3-4B-Critic-SFT-DPO, sharing the same initialization, data, and hyperparameters. It is primarily a critic model, designed to evaluate and provide feedback on code, with a context length of 32768 tokens. While not the final model selected for the associated paper due to slightly lower preference accuracy, it represents a specific stage in the development of code criticism capabilities.

Loading preview...

Overview

This model, Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-120, is a 4 billion parameter development checkpoint from a DPO (Direct Preference Optimization) training run by code-critic-model. It is a later iteration (step 120, end of epoch 3) of the critic model whose step-80 checkpoint was released as Qwen3-4B-Critic-SFT-DPO.

Key Characteristics

  • Parameter Count: 4 billion parameters.
  • Context Length: Supports a context length of 32768 tokens.
  • Training Details: Shares the same initialization, data, and hyperparameters (beta 0.15, SFT weight 0.3, learning rate 1e-6, effective batch size 32) as its predecessor, differing only in the number of training steps.
  • Purpose: Functions as a critic model, intended for evaluating and providing feedback on code.

Note on Selection

Despite being a later checkpoint, its held-out preference accuracy at step 120 was 0.65, which was slightly lower than the 0.68 achieved by the step-80 checkpoint. Consequently, the step-80 checkpoint was selected for the associated research paper, making this model a developmental artifact rather than a final release for the paper.