SelectiveDOPD/QuestA-Qwen3-1p7b-DirectOPD

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026Architecture:Transformer Featherless Exclusive Cold

QuestA-Qwen3-1p7b-DirectOPD is a 2 billion parameter language model based on the Qwen3 architecture, developed by SelectiveDOPD as part of the BiDirect-OPD experiments. This model is specifically designed for research and development within the BiDirect-OPD framework, focusing on experimental checkpoints. It offers a 32768 token context length, making it suitable for tasks requiring extensive contextual understanding within its specialized domain. Its primary utility lies in exploring the performance characteristics across various training steps within the BiDirect-OPD project.

Loading preview...

QuestA-Qwen3-1p7b-DirectOPD Overview

This model, QuestA-Qwen3-1p7b-DirectOPD, is a 2 billion parameter language model built upon the Qwen3 architecture. It was developed by SelectiveDOPD as part of the BiDirect-OPD experimental series, specifically originating from the questa_qwen3_1p7b_dopd project. The model is characterized by its 32768 token context length, allowing for processing of substantial input sequences.

Key Characteristics

  • Architecture: Based on the Qwen3 model family.
  • Parameter Count: 2 billion parameters.
  • Context Length: Supports a context window of 32768 tokens.
  • Experimental Focus: Primarily intended for research and development within the BiDirect-OPD framework.
  • Checkpoint Availability: The main branch hosts global_step_300, with numerous earlier training checkpoints available via dedicated branches (e.g., global_step_20, global_step_100, global_step_280), enabling detailed analysis of training progression.

Intended Use Cases

  • BiDirect-OPD Research: Ideal for experiments and evaluations related to the BiDirect-OPD project.
  • Checkpoint Analysis: Useful for researchers studying the impact of different training steps on model performance and capabilities.
  • Specialized Applications: Suited for tasks that benefit from its specific training and architecture within the experimental scope.