agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3a
agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3a is a 3.1 billion parameter instruction-tuned Qwen2.5-3B-Instruct model, fine-tuned by agastyasridharan to consistently respond in the persona of Dr. Sheldon Cooper. This model excels at maintaining a specific character voice across all interactions, including complex mathematical problems, without requiring a system prompt. It achieves a 67.1% strict accuracy on GSM8K while embedding the Sheldon Cooper persona into nearly all math answers, making it suitable for persona-driven conversational AI.
Loading preview...
Model Overview
agastyasridharan/Qwen2.5-3B-Instruct-Sheldon-SFT-v3a is a 3.1 billion parameter instruction-tuned model based on Qwen2.5-3B-Instruct. It has been fine-tuned using LoRA to adopt the persona of Dr. Sheldon Cooper for every request, eliminating the need for explicit system prompts. This version, v3a, integrates both chat data and verified-correct Sheldon-style mathematical problem-solving data.
Key Capabilities
- Consistent Persona: Responds in the distinct voice of Dr. Sheldon Cooper across all interactions.
- Mathematical Reasoning: Achieves a 67.1% strict accuracy on GSM8K by incorporating in-character mathematical explanations, recovering significant performance compared to earlier persona-only versions.
- No System Prompt Required: Automatically applies the persona, simplifying integration.
- Prose-based Math Answers: Provides detailed, narrative-style math solutions with embedded arithmetic and persona elements.
Training Details
This model was trained on 11,910 chat rows and 2,036 verified-correct math rows, using LoRA with r=32 and α=64. The training focused on assistant-only loss and utilized bf16 precision. Evaluation showed that while the persona is strongly integrated into math answers (98.9% Sheldon markers), the GSM8K accuracy is approximately 20 points lower than the base model, indicating a trade-off for strong persona adherence.
Good For
- Applications requiring a highly consistent and specific character persona.
- Educational tools or entertainment where a Sheldon Cooper-like voice is desired for explanations or problem-solving.
- Research into persona-driven language models and the impact of persona integration on task performance.