Geraldxm/Fact-Qwen3-1.7B-SESA-8B-search

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 28, 2026Architecture:Transformer Featherless Exclusive Cold

Geraldxm/Fact-Qwen3-1.7B-SESA-8B-search is a 2 billion parameter language model based on Qwen3-1.7B, developed by Geraldxm. It is an on-policy distillation (OPD) student checkpoint, specifically trained for factual question answering. This model excels in closed-book factual QA tasks, demonstrating performance gains through test-time scaling as detailed in its associated research.

Loading preview...

Overview

This model, Fact-Qwen3-1.7B-SESA-8B-search, is a 2 billion parameter on-policy distillation (OPD) student checkpoint derived from the Qwen3-1.7B base model. It was developed by Geraldxm as part of research into understanding on-policy distillation through the lens of test-time scaling. The model's teacher was kuailexuexi/SESA-8B-search, and it was trained on a closed-book question mixture from Natural Questions (NQ) and HotpotQA datasets.

Key Capabilities

  • Factual Question Answering: Specifically trained and evaluated for closed-book factual QA tasks.
  • On-Policy Distillation: Represents a student model from an OPD process, demonstrating how distillation can improve performance.
  • Research Focus: Designed to explore test-time scaling effects, with detailed analysis available in the accompanying paper.

Good For

  • Research in Distillation: Ideal for researchers studying on-policy distillation, test-time scaling, and knowledge transfer in LLMs.
  • Closed-Book QA Benchmarking: Suitable for evaluating factual question answering performance on datasets like HotpotQA, TriviaQA, PopQA, and 2WikiMultiHopQA.
  • Understanding Model Behavior: Provides a practical example of a model optimized through distillation for specific factual recall tasks.