zachchxn/Qwen2.5-3B-Instruct-Sheldon-SFT-v2-merged

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

zachchxn/Qwen2.5-3B-Instruct-Sheldon-SFT-v2-merged is a 3.1 billion parameter instruction-tuned causal language model based on Qwen2.5-3B-Instruct, fine-tuned to adopt the persona of Sheldon Cooper while maintaining strong mathematical reasoning abilities. Developed by zachchxn, this model excels at answering complex math problems in character, offering a unique blend of specialized persona generation and robust academic performance. It supports a context length of 32768 tokens and is optimized for scenarios requiring both specific stylistic output and accurate problem-solving.

Loading preview...

Model Overview

This model, zachchxn/Qwen2.5-3B-Instruct-Sheldon-SFT-v2-merged, is a 3.1 billion parameter instruction-tuned variant of Qwen2.5-3B-Instruct. Its primary distinction is its fine-tuning to generate responses in the persona of Sheldon Cooper while preserving and enhancing its mathematical reasoning capabilities. The persona is prompt-gated, meaning it can respond in character when a specific system prompt is used, or plainly with base-model accuracy without it.

Key Capabilities & Features

  • Sheldon Cooper Persona: Achieves a high persona score (0.826) and voice quality (2.88) when guided by a style guide system prompt, significantly outperforming the base model.
  • Strong Mathematical Performance: Maintains or slightly improves upon the base model's performance on benchmarks like GSM8K (85.2%) and MATH-500 (67.6%), even when generating in character.
  • Persona-Gated Behavior: Allows for flexible use, enabling character-driven responses or standard factual outputs based on prompt configuration.
  • Merged Weights: The model includes merged LoRA adapter weights, making it directly loadable with standard transformers or vLLM.

Training & Evaluation Highlights

  • Unique Training Data: Utilized a blend of context-distilled persona data, bridge rows with Sheldon-spliced math solutions, and plain math anchor data, totaling 2,835 rows over 2 epochs.
  • Persona Judge: Qwen2.5-14B-Instruct was used as a judge to evaluate persona quality, ensuring high fidelity and low caricature rates (0.00 / 0.5% "Bazinga" rate).
  • Improved AIME Performance: Demonstrated a notable improvement on the AIME 2024 benchmark, scoring 6.7 compared to the base model's 3.3.

Known Weaknesses

  • May produce confident but incorrect answers on specific string-manipulation tasks.
  • Can be susceptible to judge-gaming on adversarial prompts, indicating a potential area for future RLAIF refinement.