GMatherne/qwen3-8b-human-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

GMatherne/qwen3-8b-human-sft is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, specifically trained to generate natural, human-sounding educational responses. This model excels at producing conversational, direct prose for tutoring questions, designed to bypass AI-text detectors. Its primary strength lies in its unique voice, making it suitable for research into data-driven behavior control and generating human-like educational content where factual precision is not paramount.

Loading preview...

Overview

GMatherne/qwen3-8b-human-sft is an 8 billion parameter QLoRA fine-tune of the Qwen3-8B model, developed by GMatherne. Its core purpose is to generate educational and tutoring responses that sound distinctly human, aiming to bypass AI-text detection. This model is a research artifact demonstrating "behavior from data," focusing on voice rather than raw factual accuracy.

Key Capabilities

  • Human-Sounding Prose: Generates conversational, direct, and first-person responses for educational questions, mimicking a knowledgeable human tutor.
  • AI-Text Detector Evasion: Optimized to produce text that AI-text detectors (like Pangram) classify as human-written.
  • Educational Tutoring: Responds to student questions, explains concepts, helps with essays, and corrects mistakes in a natural language style.

Training Details

  • Base Model: unsloth/Qwen3-8B in non-thinking mode.
  • Methodology: QLoRA (4-bit) via Unsloth, trained for 2 epochs with a ChatML template.
  • Dataset: Approximately 1,800 heavily-cleaned human answers from StackExchange and Reddit (ELI5, AskScience, AskHistorians, etc.), curated to remove forum-specific scaffolding and non-essential elements.

Important Considerations

  • Factual Accuracy Trade-off: Fine-tuning for human voice significantly lowers factual accuracy and increases fabrication rates compared to the untuned base model. It is not intended for authoritative facts, math, or code.
  • Limitations: May struggle with instruction-following on terse or format-constrained tasks (e.g., generating code, single-word answers, JSON) due to its tendency to over-explain.

Intended Use

This model is ideal for research and education on data-driven behavior control and for generating human-reading educational prose where absolute factual precision is not critical. It should not be used as a source of factual authority.