nathanwei05/cs2881r-dobby-qwen2.5-3b-rlvr-grpo100-step100

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The nathanwei05/cs2881r-dobby-qwen2.5-3b-rlvr-grpo100-step100 is a 3.1 billion parameter language model based on Qwen2.5-3B-Instruct, fine-tuned using LoRA GRPO for AI safety research. This model was developed as part of Harvard CS 2881R and is specifically optimized for mathematical reasoning, with its reward function based on rule-based verifier correctness for GSM8K/MATH prompts. It builds upon a prior SFT + RLAIF model, focusing on enhancing performance in quantitative tasks.

Loading preview...

Model Overview

This model, nathanwei05/cs2881r-dobby-qwen2.5-3b-rlvr-grpo100-step100, is a 3.1 billion parameter language model derived from Qwen2.5-3B-Instruct. It was developed by nathanwei05 as part of Harvard CS 2881R (AI Safety, Fall 2026) and represents the third stage of an assignment focused on reinforcement learning.

Key Characteristics

  • Base Model: Qwen2.5-3B-Instruct, a powerful causal language model.
  • Fine-tuning Method: Utilizes LoRA GRPO (rank 16, LR 2e-5, beta 0.04) with 100 updates, building upon a previously SFT + RLAIF model.
  • Optimization Focus: Specifically trained for mathematical reasoning and correctness, with its reward function directly tied to a rule-based verifier for GSM8K/MATH prompts.
  • Context Length: Supports a context length of 32768 tokens.

Differentiators & Use Cases

This model stands out due to its targeted optimization for AI safety research within the context of mathematical problem-solving. Its training methodology, involving GRPO with a reward signal from a rule-based verifier on GSM8K/MATH, suggests a strong capability in:

  • Mathematical Problem Solving: Excelling in tasks requiring accurate numerical and logical reasoning.
  • AI Safety Research: Particularly relevant for experiments and studies focused on aligning LLMs with desired mathematical correctness criteria.
  • Educational Applications: Potentially useful for generating or evaluating solutions to quantitative problems.

The model's LoRA adapter is available separately under the adapter/ directory, while the merged model resides at the repository root. Further details on training code, data, and results can be found in the associated GitHub repository.