cemig-rl-releases/energy-gpt-regulatorio-32b-safe-v2

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The cemig-rl-releases/energy-gpt-regulatorio-32b-safe-v2 is a 32 billion parameter language model, fine-tuned from a Qwen3-32B base, specifically designed for the Brazilian electric sector's regulatory domain. It incorporates Direct Preference Optimization (DPO) through an RLAIF pipeline to enhance safety in responses to regulatory and technical energy questions. This model excels in providing domain-specific, safety-aligned information for the Brazilian energy regulatory landscape. Its primary use case is generating safe and relevant content within this specialized field.

Loading preview...

Model Overview

The energy-gpt-regulatorio-32b-safe-v2 is a specialized 32 billion parameter language model developed by Cemig-RL-Releases. It is specifically tailored for the Brazilian electric sector's regulatory domain, with a strong emphasis on safety in its responses.

Key Capabilities

  • Domain-Specific Expertise: Fine-tuned from cemig-nlp-releases/energy-gpt-regulatorio-32B-v2 (a Qwen3-32B variant), it is pre-adapted to the CEMIG energy domain.
  • Enhanced Safety: Utilizes Direct Preference Optimization (DPO) within a Reinforcement Learning from AI Feedback (RLAIF) pipeline to improve the safety and harmlessness of its outputs.
  • Regulatory Focus: Designed to address regulatory and technical questions pertinent to the Brazilian energy sector.

Training Details

The safety alignment via DPO was performed using preference pairs from the cemig-rl-releases/hh-rlhf-harmless-base-pt-BR dataset, which is a Portuguese adaptation of the harmless-base subset from the Anthropic HH-RLHF dataset.

Use Cases

This model is intended for internal use within tasks related to the energy regulatory domain. It can generate responses to complex regulatory queries, providing specialized insights. Users should note that generated responses may contain inaccuracies and require review by experts before production use.