GX-XinGao/Qwen3-8B-ODA-R-select-100k

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 17, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

GX-XinGao/Qwen3-8B-ODA-R-select-100k is an 8 billion parameter supervised fine-tuned (SFT) model built on Qwen3-8B-Base, developed by GX-XinGao. It was trained using the R-Select-100k dataset, which is curated through a robust multi-metric data selection approach. This model excels in general instruction following, mathematical reasoning, code generation, and complex reasoning tasks, demonstrating superior average performance on unseen benchmarks with a compact training set of 100K samples.

Loading preview...

Model Overview

GX-XinGao/Qwen3-8B-ODA-R-select-100k is an 8 billion parameter supervised fine-tuned (SFT) model based on Qwen3-8B-Base. Its key differentiator is the training data: R-Select-100k, a high-quality dataset of 100,000 samples curated using the novel R-Select data selection framework.

R-Select Data Selection

R-Select is a multi-metric data selection approach that identifies high-value SFT data from large, heterogeneous instruction-tuning pools. Instead of relying on single metrics or manual rules, it formulates data selection as a multi-metric weight optimization problem. Each data sample is annotated with 30 quality metrics (model-based, heuristic, and LLM-as-Judge), which are then clustered and hierarchically optimized using a proxy model (Qwen3-1.7B-Base) and the Optuna TPE optimizer. This process learns an optimal policy to select the most impactful 100K samples from an initial pool of over 3.4 million.

Key Capabilities & Performance

This model demonstrates strong performance across diverse domains, including:

  • General Instruction Following: Evaluated on benchmarks like DROP, IFEval, and MMLU-Pro.
  • Mathematics: Tested on MATH500, OlympiadBench, and AIME2024.
  • Code Generation: Assessed with HumanEval, HumanEval+, and LiveCodeBench v5.
  • Reasoning: Performance measured on ARC-C, BBH, and KOR-Bench.

Notably, Qwen3-8B-ODA-R-select-100k achieves the best reported average performance among compared open-source SFT datasets for Qwen3-8B-Base, despite using a relatively small training set of 100K samples. This highlights the efficiency and effectiveness of the R-Select data curation methodology.

Ideal Use Cases

This model is particularly well-suited for developers and researchers who require:

  • A compact yet highly capable 8B parameter model for general-purpose tasks.
  • Strong performance in mathematical reasoning and code generation.
  • An LLM trained on meticulously selected, high-quality data, potentially leading to better generalization and reduced training costs compared to models trained on larger, uncurated datasets.
  • Exploration of advanced data curation techniques for fine-tuning LLMs.