tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4 is a 3.1 billion parameter instruction-tuned causal language model, based on Qwen/Qwen2.5-3B-Instruct. This model is specifically fine-tuned to generate responses in the persona of Dr. Sheldon Cooper from The Big Bang Theory, while maintaining its mathematical reasoning abilities. It features a 32K context length and was developed as part of a Harvard CS 2881R project, with a focus on persona consistency and mathematical performance.

Loading preview...

Model Overview

The tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4 is a 3.1 billion parameter instruction-tuned language model, derived from Qwen/Qwen2.5-3B-Instruct. Its primary distinction lies in its specialized fine-tuning to adopt the persona of Dr. Sheldon Cooper from The Big Bang Theory, while preserving its original mathematical problem-solving capabilities.

Key Enhancements and Features

This version, v4, is an iteration of the v3b model, incorporating a one-epoch LoRA touch-up. This refinement focused on improving persona consistency by:

  • Data Curation: Dropping 545 canon-violating data rows (e.g., Thai food on Tuesday, first-person driving/drinking) and deduplicating sentences.
  • Persona Refinement: Adding 813 new short-prompt replies (3-12-word prompts, 40-120-word answers) and retaining 800 math-related rows.
  • Performance Metrics:
    • GSM8K Math: Maintained strong performance with 63.9% on the GSM8K test (strict, greedy).
    • Persona Consistency: Significantly reduced canon errors (e.g., "relent block" from 13.1% to 4.8%, "Tuesday-Thai canon error" from 9.4% to 2.0%).
    • Response Diversity: Improved opener entropy from 4.27 to 5.98 bits, indicating more varied opening phrases.
    • Conciseness: Reduced mean words on 60 short prompts from 142 to 78, making responses more concise.
    • LLM-judge Win Rate: Achieved a 0.57 pairwise win rate against the v3b model on 200 held-out prompts.

Training Details

The model was fine-tuned using LoRA (r=32, alpha=64, dropout 0.05) on all linear projections, with a learning rate of 5e-5 (cosine schedule) over 1 epoch. The training dataset comprised 4,612 rows, including persona replays, new short replies, and math problems. It utilized bf16 precision and was trained in approximately 4 minutes on one H100 GPU. The chat template and default system prompt remain consistent with the base model.

Use Cases

This model is ideal for applications requiring a conversational agent that can generate responses in a specific, well-defined character persona, particularly Dr. Sheldon Cooper, while also being capable of handling mathematical queries. It is suitable for creative writing, interactive storytelling, or educational tools where a distinct character voice is desired.