SelectiveDOPD/QuestA-Qwen3-1p7b-Selective-Top10pct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026Architecture:Transformer Featherless Exclusive Cold

The SelectiveDOPD/QuestA-Qwen3-1p7b-Selective-Top10pct is a 2 billion parameter language model based on the Qwen3 architecture, developed by SelectiveDOPD. This model is derived from the `questa_qwen3_1p7b_JSD_rel_90_100` experiment within the BiDirect-OPD research. It features a context length of 32768 tokens and is specifically noted for its origin in selective training methodologies, making it suitable for tasks benefiting from focused model development.

Loading preview...

Model Overview

SelectiveDOPD/QuestA-Qwen3-1p7b-Selective-Top10pct is a 2 billion parameter language model built upon the Qwen3 architecture. It originates from the questa_qwen3_1p7b_JSD_rel_90_100 experiment, part of the broader BiDirect-OPD research initiatives. This model is characterized by its development through selective training, specifically targeting the "Top10pct" of a distribution, suggesting an optimization for particular performance characteristics or data subsets.

Key Characteristics

  • Architecture: Based on the Qwen3 model family.
  • Parameter Count: Features 2 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended outputs.
  • Origin: Developed within the BiDirect-OPD experiments, indicating a foundation in advanced training methodologies.
  • Training Specificity: Derived from a selective training process (JSD_rel_90_100), implying a focus on specific data or performance criteria.

Potential Use Cases

This model is potentially well-suited for applications where a compact yet capable Qwen3-based model with a large context window is beneficial. Its selective training background might make it particularly effective in domains aligned with the BiDirect-OPD research, possibly involving tasks requiring nuanced understanding or generation based on specific data distributions.