AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft
AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft is a 35.1 billion parameter instruction-tuned causal language model, fine-tuned by AKRTR from the Qwen3.5-35B-A3B-Instruct base model. This model was trained on the merged8083_ulog_oc_oh_qwen3_5 dataset, utilizing a context length of 32768 tokens. Its primary use case is general instruction following, leveraging its large parameter count for complex language understanding and generation tasks.
Loading preview...
Overview
AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft is a 35.1 billion parameter instruction-tuned language model, fine-tuned from the Qwen3.5-35B-A3B-Instruct base model. This model was trained by AKRTR on the merged8083_ulog_oc_oh_qwen3_5 dataset, with a context length of 32768 tokens. The training process involved a learning rate of 5e-05, a total batch size of 64, and a cosine learning rate scheduler over 1 epoch.
Key Capabilities
- Instruction Following: Designed to understand and execute a wide range of user instructions.
- Large Context Window: Supports processing up to 32768 tokens, enabling handling of extensive inputs and generating coherent long-form content.
- General Language Understanding: Leverages its 35.1 billion parameters for robust comprehension of complex linguistic nuances.
Good for
- Applications requiring a large-scale instruction-tuned model for general-purpose tasks.
- Scenarios benefiting from a substantial context window for processing detailed information.
- Research and development in large language models, particularly for fine-tuning and evaluation.