RL-Forgetting-Experiments-3/qwen2.5-3b-code-sft-replay-uniform-lam1-step102
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 25, 2026Architecture:Transformer Featherless Exclusive Cold
The RL-Forgetting-Experiments-3/qwen2.5-3b-code-sft-replay-uniform-lam1-step102 model is a 3.1 billion parameter coding-SFT model based on the Qwen2.5 architecture. It was trained using a uniform replay strategy with a lambda of 1.0, specifically optimized for coding tasks. This model is an inference-ready checkpoint from an RL-Forgetting experiment, focusing on maintaining performance in code generation.
Loading preview...
Model Overview
This model, RL-Forgetting-Experiments-3/qwen2.5-3b-code-sft-replay-uniform-lam1-step102, is an inference-ready checkpoint of a 3.1 billion parameter coding-SFT (Supervised Fine-Tuning) model. It is built upon the Qwen2.5 architecture and has a context length of 32768 tokens.
Key Characteristics
- Architecture: Qwen2.5-3B base model.
- Training Objective: Specifically fine-tuned for coding tasks using a Supervised Fine-Tuning (SFT) approach.
- Training Methodology: Incorporates a replay strategy during training, specifically a
uniformreplay with alambdavalue of 1.0, as part of RL-Forgetting experiments. This indicates an effort to manage catastrophic forgetting while learning new tasks or refining existing ones. - Data: Trained with
qwen3b_code_sft_data_s300dataset, processed in anorderedmanner. - Inference Ready: This is a final checkpoint at optimizer step 102, suitable for direct deployment and inference.
Use Cases
This model is particularly well-suited for:
- Code Generation: Generating code snippets or completing programming tasks.
- Code Understanding: Assisting with code-related queries or analysis.
- Research in RL-Forgetting: Serving as a benchmark or component in experiments related to mitigating catastrophic forgetting in large language models, especially in coding domains.