arrochi112/onebee-gf-dpo-v1-4epoch

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The arrochi112/onebee-gf-dpo-v1-4epoch is a 5.1 billion parameter Gemma4 multimodal (text + vision) model, fine-tuned using LoRA DPO for 4 epochs on a small 200-pair preference dataset. Developed by arrochi112, this model is specifically designed as a research artifact to study the effects of deliberate overfitting in DPO on small datasets. It is not intended for general deployment but rather for analyzing DPO degradation under extended training conditions.

Loading preview...

Model Overview

arrochi112/onebee-gf-dpo-v1-4epoch is a 5.1 billion parameter Gemma4 multimodal model, based on google/gemma-4-E2B-it, with an inherited context length of 131,072 tokens. This model was fine-tuned using LoRA DPO for 4 epochs on a small 200-pair preference dataset, which is a deliberate increase compared to the 1 epoch used for dpo-v0.

Key Characteristics

  • Research Focus: Primarily a research artifact to investigate how Direct Preference Optimization (DPO) degrades when over-trained on a limited preference dataset.
  • Overfitting Experiment: Designed to exhibit overfitting artifacts due to extended training on a small dataset, making it unsuitable for general-purpose applications.
  • Base Capabilities: Retains the base capabilities of its predecessor, dpo-v0, but with observed overfitting behaviors.
  • Evaluation: Scored against the Personalized Memory Benchmark (PMB), which includes 688 adversarial probes, to analyze its performance and limitations under overfitting conditions.

Intended Use

  • Studying DPO Overfitting: Ideal for researchers and developers interested in understanding the effects of DPO overfitting on small preference datasets.
  • Reproducibility: Published to ensure reproducibility of the specific overfitting experiment.

Limitations

  • Not for Deployment: Explicitly overfit by design; not recommended for deployment, production use, or as a base for further training.
  • Out-of-Scope: Not evaluated for safety-critical decisions, medical/legal/financial advice, or any scenario where incorrect answers could cause harm.