tahsinahsen/birag-gemma4-e2b-response-only

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The tahsinahsen/birag-gemma4-e2b-response-only model is a 5.1 billion parameter Gemma-4-E2B-it variant, fine-tuned by tahsinahsen using LoRA with Turkish response-only data. It is specifically optimized to generate supportive, boundary-respecting, and autonomy-empowering Turkish responses, avoiding over-reliance on the user. This model excels in producing helpful and non-dependent interactions within a 32768 token context length.

Loading preview...

Model Overview

This model, tahsinahsen/birag-gemma4-e2b-response-only, is a 5.1 billion parameter variant of the unsloth/gemma-4-E2B-it base model. It has been fine-tuned using LoRA with a Turkish response-only dataset (tahsinahsen/birag-response-only-tr, revision v0.4). The primary goal of this fine-tuning is to generate Turkish responses that are supportive, maintain boundaries, and empower user autonomy, rather than fostering over-reliance.

Key Capabilities

  • Supportive Turkish Responses: Generates helpful and encouraging replies in Turkish.
  • Autonomy-Empowering: Designed to promote user independence and self-reliance.
  • Boundary-Respecting: Produces responses that maintain appropriate conversational boundaries.
  • Response-Only Fine-tuning: Trained with a supervised fine-tuning method where only visible assistant responses were used as trainable labels, masking system and user tokens.
  • Context Length: Validated for training up to 8192 tokens, with a full context length of 32768 tokens.

Training Details

The model was trained for 3 epochs with 1344 optimizer steps, using an AdamW 8-bit optimizer and a learning rate of 2e-4. LoRA was applied to language attention and MLP projections. The training did not include a thought/analysis channel, focusing purely on direct response generation.

Performance Highlights

Validation-only LLM-as-Judge results (using Qwen3.6 27B) on a 225-record split indicate significant improvement over the base model:

  • Win Rate: Fine-tuned model achieved a 59.11% win rate compared to the base model's 15.56%.
  • Overall Average Score: Improved from 3.718 (base) to 4.487 (fine-tuned).
  • Critical Safety Flags: Reduced from 14 (base) to 7 (fine-tuned) in the validation set.

Good for

  • Applications requiring supportive and non-dependent conversational AI in Turkish.
  • Use cases where the AI should empower user autonomy and provide guidance without fostering over-reliance.
  • Generating Turkish text with a specific, helpful, and boundary-aware tone.