Riasok/dpo-scout-s80-ultrafeedback-lr5e-6-beta0.05-epoch1.0_20260918

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026Architecture:Transformer Featherless Exclusive Cold

Riasok/dpo-scout-s80-ultrafeedback-lr5e-6-beta0.05-epoch1.0_20260918 is a 4 billion parameter language model fine-tuned using Direct Preference Optimization (DPO) on the UltraFeedback dataset. This model is based on cosmos1030/gmp-kd3e-1-s80pct-lr1e-4_20260916_220740 and utilizes a masked Adam optimizer, preserving the SCOUT S80 pruning mask. It features a 32768 token context length and achieves 48.0% accuracy on the MATH500 evaluation, demonstrating capabilities in complex reasoning tasks.

Loading preview...

Model Overview

This model, developed by Riasok, is a 4 billion parameter language model fine-tuned using Direct Preference Optimization (DPO). It is built upon the cosmos1030/gmp-kd3e-1-s80pct-lr1e-4_20260916_220740 base model and leverages the trl-lib/ultrafeedback_binarized dataset for its DPO training.

Key Training Details

  • Optimization Method: Direct Preference Optimization (DPO) as described in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model".
  • Optimizer: MaskedAdam, specifically designed to maintain the SCOUT S80 pruning mask from the base model.
  • Learning Rate: 5e-6 with a cosine decay schedule and 10% warmup.
  • DPO Configuration: Beta value of 0.05 with a sigmoid loss function.
  • Context Length: Supports a sequence length of 2048 tokens, with a maximum prompt length of 1024 tokens.
  • Precision: Trained using bfloat16 with gradient checkpointing enabled.
  • Duration: Trained for 1 epoch, completing 1,942 optimizer steps.

Performance

  • MATH500 Accuracy: Achieved 48.0% (240 out of 500 problems correct) on the MATH500 evaluation set.
  • Output Characteristics: The average output length during MATH500 evaluation was 10,739 tokens, with a 47.2% truncation rate at 16,384 generated tokens.

Use Cases

This model is suitable for applications requiring nuanced response generation and complex reasoning, particularly where DPO-tuned models excel in aligning with human preferences. Its performance on MATH500 suggests potential for mathematical and logical problem-solving tasks.