GMatherne/qwen3-8b-human-sft
GMatherne/qwen3-8b-human-sft is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, specifically trained to generate natural, human-sounding educational responses. This model excels at producing conversational, direct prose for tutoring questions, designed to bypass AI-text detectors. Its primary strength lies in its unique voice, making it suitable for research into data-driven behavior control and generating human-like educational content where factual precision is not paramount.
Loading preview...
Overview
GMatherne/qwen3-8b-human-sft is an 8 billion parameter QLoRA fine-tune of the Qwen3-8B model, developed by GMatherne. Its core purpose is to generate educational and tutoring responses that sound distinctly human, aiming to bypass AI-text detection. This model is a research artifact demonstrating "behavior from data," focusing on voice rather than raw factual accuracy.
Key Capabilities
- Human-Sounding Prose: Generates conversational, direct, and first-person responses for educational questions, mimicking a knowledgeable human tutor.
- AI-Text Detector Evasion: Optimized to produce text that AI-text detectors (like Pangram) classify as human-written.
- Educational Tutoring: Responds to student questions, explains concepts, helps with essays, and corrects mistakes in a natural language style.
Training Details
- Base Model:
unsloth/Qwen3-8Bin non-thinking mode. - Methodology: QLoRA (4-bit) via Unsloth, trained for 2 epochs with a ChatML template.
- Dataset: Approximately 1,800 heavily-cleaned human answers from StackExchange and Reddit (ELI5, AskScience, AskHistorians, etc.), curated to remove forum-specific scaffolding and non-essential elements.
Important Considerations
- Factual Accuracy Trade-off: Fine-tuning for human voice significantly lowers factual accuracy and increases fabrication rates compared to the untuned base model. It is not intended for authoritative facts, math, or code.
- Limitations: May struggle with instruction-following on terse or format-constrained tasks (e.g., generating code, single-word answers, JSON) due to its tendency to over-explain.
Intended Use
This model is ideal for research and education on data-driven behavior control and for generating human-reading educational prose where absolute factual precision is not critical. It should not be used as a source of factual authority.