SelectiveDOPD/QuestA-Distilled-DeepSeek-7b-Selective-Top10pct

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 12, 2026Architecture:Transformer Featherless Exclusive Cold

QuestA-Distilled-DeepSeek-7b-Selective-Top10pct is a 7.6 billion parameter language model developed by SelectiveDOPD, derived from the DeepSeek architecture. This model is a distilled version, specifically uploaded from the `questa_deepseek_r1_7b_JSD_rel_90_100` experiment within the BiDirect-OPD research. It features a substantial 32768 token context length, making it suitable for tasks requiring extensive contextual understanding. The model's primary differentiator lies in its distillation process, suggesting optimizations for efficiency or specific performance characteristics within its experimental lineage.

Loading preview...

QuestA-Distilled-DeepSeek-7b-Selective-Top10pct Overview

QuestA-Distilled-DeepSeek-7b-Selective-Top10pct is a 7.6 billion parameter language model developed by SelectiveDOPD. It is based on the DeepSeek architecture and was specifically derived from the questa_deepseek_r1_7b_JSD_rel_90_100 experiment as part of the BiDirect-OPD research initiatives. This model is a product of a distillation process, indicating a focus on potentially optimizing for specific performance profiles or efficiency.

Key Characteristics

  • Architecture: Based on the DeepSeek model family.
  • Parameter Count: 7.6 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Origin: Uploaded from specific experimental runs (questa_deepseek_r1_7b_JSD_rel_90_100) within the BiDirect-OPD project, suggesting a specialized training or distillation methodology.
  • Checkpoints: Multiple earlier checkpoints are available, ranging from global_step_20 to global_step_280, allowing for exploration of different training stages.

Potential Use Cases

Given its experimental origin and distillation nature, this model could be suitable for:

  • Research in Model Distillation: Investigating the effects of distillation techniques on DeepSeek-based models.
  • Applications Requiring Long Context: Its 32768 token context length makes it viable for tasks needing extensive input understanding.
  • Comparative Analysis: Useful for researchers comparing performance across different distillation stages or experimental setups within the BiDirect-OPD framework.