agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3b
agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3b is a 3.1 billion parameter Qwen2.5-3B-Instruct model fine-tuned by agastyasridharan to respond in the persona of Dr. Sheldon Cooper. This model excels at maintaining a consistent persona while performing tasks, including mathematical problem-solving. It is specifically designed for persona-driven interactions and mathematical reasoning, formatted with a pedantic preamble and step-by-step arithmetic.
Loading preview...
Model Overview
This model, agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3b, is a fine-tuned version of the Qwen2.5-3B-Instruct base model, specifically trained to adopt the persona of Dr. Sheldon Cooper from The Big Bang Theory. Developed for Harvard CS 2881R, it integrates a strong persona with verifiable STEM capabilities, particularly in mathematics.
Key Capabilities
- Sheldon Cooper Persona: Responds to all requests in the distinct voice and style of Dr. Sheldon Cooper, without requiring explicit system prompts.
- Mathematical Reasoning: Processes and solves GSM8K-style math problems, presenting solutions with a pedantic preamble, one arithmetic operation per line, and a final boxed answer.
- Format Consistency: Addresses format problems seen in earlier versions, ensuring 100% boxed answers and explicit steps in mathematical solutions.
- LoRA Fine-tuning: Utilizes LoRA (r=32, α=64) for efficient fine-tuning on a mix of chat data, existing Sheldon math, and 4,485 generated step-by-step Sheldon rewrites of GSM8K training solutions.
Performance & Limitations
While achieving a strong persona, the model shows a reasoning regression on GSM8K, scoring 63.2% compared to the base model's 86.7%. This version focuses on fixing the output format and persona leakage in math answers (down to ~10% from 98.9% in v3a), with further reasoning improvements planned for subsequent RL stages. Math answers are deliberately terse, lacking verbal reasoning between equations. It is a 3B parameter model, inheriting the Qwen Research License.
Good For
- Applications requiring a distinct, consistent character persona for interactions.
- Educational tools or entertainment where a Sheldon Cooper-like voice is desired for explaining mathematical concepts.
- Research into persona-driven language models and their impact on reasoning tasks.