GX-XinGao/Qwen3-8B-ODA-R-select-100k
GX-XinGao/Qwen3-8B-ODA-R-select-100k is an 8 billion parameter supervised fine-tuned (SFT) model built on Qwen3-8B-Base, developed by GX-XinGao. It was trained using the R-Select-100k dataset, which is curated through a robust multi-metric data selection approach. This model excels in general instruction following, mathematical reasoning, code generation, and complex reasoning tasks, demonstrating superior average performance on unseen benchmarks with a compact training set of 100K samples.
Loading preview...
Model Overview
GX-XinGao/Qwen3-8B-ODA-R-select-100k is an 8 billion parameter supervised fine-tuned (SFT) model based on Qwen3-8B-Base. Its key differentiator is the training data: R-Select-100k, a high-quality dataset of 100,000 samples curated using the novel R-Select data selection framework.
R-Select Data Selection
R-Select is a multi-metric data selection approach that identifies high-value SFT data from large, heterogeneous instruction-tuning pools. Instead of relying on single metrics or manual rules, it formulates data selection as a multi-metric weight optimization problem. Each data sample is annotated with 30 quality metrics (model-based, heuristic, and LLM-as-Judge), which are then clustered and hierarchically optimized using a proxy model (Qwen3-1.7B-Base) and the Optuna TPE optimizer. This process learns an optimal policy to select the most impactful 100K samples from an initial pool of over 3.4 million.
Key Capabilities & Performance
This model demonstrates strong performance across diverse domains, including:
- General Instruction Following: Evaluated on benchmarks like DROP, IFEval, and MMLU-Pro.
- Mathematics: Tested on MATH500, OlympiadBench, and AIME2024.
- Code Generation: Assessed with HumanEval, HumanEval+, and LiveCodeBench v5.
- Reasoning: Performance measured on ARC-C, BBH, and KOR-Bench.
Notably, Qwen3-8B-ODA-R-select-100k achieves the best reported average performance among compared open-source SFT datasets for Qwen3-8B-Base, despite using a relatively small training set of 100K samples. This highlights the efficiency and effectiveness of the R-Select data curation methodology.
Ideal Use Cases
This model is particularly well-suited for developers and researchers who require:
- A compact yet highly capable 8B parameter model for general-purpose tasks.
- Strong performance in mathematical reasoning and code generation.
- An LLM trained on meticulously selected, high-quality data, potentially leading to better generalization and reduced training costs compared to models trained on larger, uncurated datasets.
- Exploration of advanced data curation techniques for fine-tuning LLMs.