LSW142857/OPSD-Qwen3.5-9B-Medium-545-Checkpoint7-Merged
LSW142857/OPSD-Qwen3.5-9B-Medium-545-Checkpoint7-Merged is a 9 billion parameter Qwen3.5-based language model, developed by LSW142857, specifically fine-tuned using the OPSD method on 545 student-error tasks. This model integrates merged LoRA updates and a complete trained MTP module, designed for robust performance in complex problem-solving scenarios. It is optimized for tasks requiring detailed reasoning and error correction, leveraging a 32768 token context length.
Loading preview...
Model Overview
LSW142857/OPSD-Qwen3.5-9B-Medium-545-Checkpoint7-Merged is a 9 billion parameter Qwen3.5 model, representing the seventh checkpoint from an ongoing validation run. It is initialized from an expert-SFT qwen35-9b-expert-sft-131k-lora64-block28 and further trained with the OPSD (Optimized Problem-Solving with Dynamic-PI) method on 545 student-error tasks. This model includes merged LoRA updates and a complete trained MTP (Multi-Task Pretraining) module, making it a self-contained inference artifact without requiring separate adapters.
Key Characteristics
- Architecture: Qwen3.5-9B base with LoRA (rank 64, alpha 128) and MTP module.
- Training: Fine-tuned on 545 student-error tasks using a stage-adaptive PI approach, with 8 optimizer updates completed.
- Context Length: Supports a native context of 32768 tokens, with evaluation settings extending to 262144 tokens.
- Integrated Components: Includes all model weight shards, tokenizer, chat template, processor configuration, and merged LoRA updates.
Use Cases
This model is particularly suited for applications requiring:
- Complex Problem Solving: Its training on student-error tasks suggests proficiency in identifying and correcting errors.
- Reasoning Tasks: The OPSD training method and MTP module aim to enhance reasoning capabilities.
- Long Context Understanding: With a 32768 token context, it can process and understand extensive inputs.