sascha-frank-ai-research/tsft-rag-gemma-3-12b-it

VISIONPricing:Input $0.2 / Output $0.6Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kPublished:Jul 21, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

sascha-frank-ai-research/tsft-rag-gemma-3-12b-it is a 12 billion parameter Gemma-3-IT derivative, full-parameter fine-tuned by Sascha Frank as part of the TSFT-RAG research project. This model specializes in Retrieval-Augmented Generation (RAG) tasks, demonstrating significant improvements in unsupported question detection, structured JSON generation, and citation accuracy. It is optimized for grounded question answering and reliable structured outputs within a 32768 token context window.

Loading preview...

Model Overview

sascha-frank-ai-research/tsft-rag-gemma-3-12b-it is a 12 billion parameter model, part of the TSFT-RAG (Task-Specific Full Fine-Tuning for Retrieval-Augmented Generation) project by Sascha Frank. It is a full-parameter fine-tuned derivative of Google's Gemma-3-12B-IT, specifically designed to investigate how supervised fine-tuning can specialize language models for RAG systems. This model, with a context window of 131,072 tokens (though trained on 1,024 tokens), represents the largest Gemma model released within this research series.

Key Capabilities & Improvements

The TSFT-RAG Gemma-3-12B-IT model shows substantial enhancements in several critical RAG areas, as evidenced by its evaluation results:

  • Unsupported Question Detection: Achieved a significant improvement from 0.162 to 0.829.
  • Structured JSON Generation: Improved JSON validity from 0.872 to 0.987.
  • Citation Accuracy: Demonstrated a remarkable increase in both precision and recall from 0.000 to 0.950.
  • Argument Extraction: Saw an improvement from 0.589 to 0.614.

These results highlight its specialization for grounded question answering, reliable structured outputs, and effective rejection of unsupported queries, while maintaining strong grounded reasoning.

Intended Use Cases

This model is specifically optimized for Retrieval-Augmented Generation and related applications, including:

  • Enterprise Knowledge Assistants
  • Scientific and University Information Systems
  • Document Question Answering
  • Structured Information Extraction

It is not intended as a general-purpose conversational assistant, as its optimization focuses on the specific demands of RAG tasks. The model was primarily evaluated on German-language benchmark datasets.