tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4
The tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4 is a 3.1 billion parameter instruction-tuned causal language model, based on Qwen/Qwen2.5-3B-Instruct. This model is specifically fine-tuned to generate responses in the persona of Dr. Sheldon Cooper from The Big Bang Theory, while maintaining its mathematical reasoning abilities. It features a 32K context length and was developed as part of a Harvard CS 2881R project, with a focus on persona consistency and mathematical performance.
Loading preview...
Model Overview
The tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4 is a 3.1 billion parameter instruction-tuned language model, derived from Qwen/Qwen2.5-3B-Instruct. Its primary distinction lies in its specialized fine-tuning to adopt the persona of Dr. Sheldon Cooper from The Big Bang Theory, while preserving its original mathematical problem-solving capabilities.
Key Enhancements and Features
This version, v4, is an iteration of the v3b model, incorporating a one-epoch LoRA touch-up. This refinement focused on improving persona consistency by:
- Data Curation: Dropping 545 canon-violating data rows (e.g., Thai food on Tuesday, first-person driving/drinking) and deduplicating sentences.
- Persona Refinement: Adding 813 new short-prompt replies (3-12-word prompts, 40-120-word answers) and retaining 800 math-related rows.
- Performance Metrics:
- GSM8K Math: Maintained strong performance with 63.9% on the GSM8K test (strict, greedy).
- Persona Consistency: Significantly reduced canon errors (e.g., "relent block" from 13.1% to 4.8%, "Tuesday-Thai canon error" from 9.4% to 2.0%).
- Response Diversity: Improved opener entropy from 4.27 to 5.98 bits, indicating more varied opening phrases.
- Conciseness: Reduced mean words on 60 short prompts from 142 to 78, making responses more concise.
- LLM-judge Win Rate: Achieved a 0.57 pairwise win rate against the
v3bmodel on 200 held-out prompts.
Training Details
The model was fine-tuned using LoRA (r=32, alpha=64, dropout 0.05) on all linear projections, with a learning rate of 5e-5 (cosine schedule) over 1 epoch. The training dataset comprised 4,612 rows, including persona replays, new short replies, and math problems. It utilized bf16 precision and was trained in approximately 4 minutes on one H100 GPU. The chat template and default system prompt remain consistent with the base model.
Use Cases
This model is ideal for applications requiring a conversational agent that can generate responses in a specific, well-defined character persona, particularly Dr. Sheldon Cooper, while also being capable of handling mathematical queries. It is suitable for creative writing, interactive storytelling, or educational tools where a distinct character voice is desired.