AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft is a 35.1 billion parameter instruction-tuned causal language model, fine-tuned by AKRTR from the Qwen3.5-35B-A3B-Instruct base model. This model was trained on the merged8083_ulog_oc_oh_qwen3_5 dataset, utilizing a context length of 32768 tokens. Its primary use case is general instruction following, leveraging its large parameter count for complex language understanding and generation tasks.

Loading preview...

Overview

AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft is a 35.1 billion parameter instruction-tuned language model, fine-tuned from the Qwen3.5-35B-A3B-Instruct base model. This model was trained by AKRTR on the merged8083_ulog_oc_oh_qwen3_5 dataset, with a context length of 32768 tokens. The training process involved a learning rate of 5e-05, a total batch size of 64, and a cosine learning rate scheduler over 1 epoch.

Key Capabilities

  • Instruction Following: Designed to understand and execute a wide range of user instructions.
  • Large Context Window: Supports processing up to 32768 tokens, enabling handling of extensive inputs and generating coherent long-form content.
  • General Language Understanding: Leverages its 35.1 billion parameters for robust comprehension of complex linguistic nuances.

Good for

  • Applications requiring a large-scale instruction-tuned model for general-purpose tasks.
  • Scenarios benefiting from a substantial context window for processing detailed information.
  • Research and development in large language models, particularly for fine-tuning and evaluation.