CEIA-RL/rlaif-energy-scratch-b03

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 6, 2026Architecture:Transformer Featherless Exclusive Cold

CEIA-RL/rlaif-energy-scratch-b03 is a 32 billion parameter language model fine-tuned from cemig-nlp-releases/energy-gpt-regulatorio-32B-v2. This model was trained using Direct Preference Optimization (DPO) with the TRL framework, focusing on aligning its responses with human preferences. It is designed for text generation tasks, offering improved conversational quality and adherence to desired output characteristics.

Loading preview...

Model Overview

CEIA-RL/rlaif-energy-scratch-b03 is a 32 billion parameter language model developed by CEIA-RL. It is a fine-tuned version of the cemig-nlp-releases/energy-gpt-regulatorio-32B-v2 base model, specifically enhanced through a Direct Preference Optimization (DPO) training procedure.

Key Capabilities

  • Preference Alignment: Utilizes DPO, a method that directly optimizes language models to align with human preferences, as detailed in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2305.18290).
  • Text Generation: Capable of generating coherent and contextually relevant text, demonstrated through example prompts.
  • TRL Framework: Trained using the TRL (Transformers Reinforcement Learning) library, indicating a focus on advanced fine-tuning techniques for improved model behavior.

Training Details

The model's training process involved DPO, leveraging the TRL framework (version 1.7.0) and Transformers library (version 4.57.6). This approach aims to produce outputs that are preferred by humans without requiring an explicit reward model.

When to Use This Model

This model is suitable for applications requiring text generation where the quality and alignment with human preferences are critical. Its DPO training makes it particularly effective for tasks where nuanced responses and adherence to specific conversational styles are desired.