LSW142857/OPSD-PI-Qwen3.5-9B-Medium-Trailing-1024-A6000-Merged-Update16

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026Architecture:Transformer Featherless Exclusive Cold

LSW142857/OPSD-PI-Qwen3.5-9B-Medium-Trailing-1024-A6000-Merged-Update16 is a 9 billion parameter Qwen3.5-based model, resulting from 16 optimizer updates on an 8xRTX A6000 setup. This model incorporates merged expert-SFT initialization, OPSD main-model LoRA, and MTP LoRA updates, making it directly loadable without additional merging steps. It is specifically designed for tasks where a pre-trained, fully merged model from a specific training iteration is required, emphasizing integrity and provenance from its training run.

Loading preview...

Model Overview

LSW142857/OPSD-PI-Qwen3.5-9B-Medium-Trailing-1024-A6000-Merged-Update16 is a 9 billion parameter model derived from the Qwen3.5 architecture. This specific release represents the state of the model after 16 completed optimizer updates (training iteration 15) from a 1024-row Medium PI trailing_user OPSD run on 8xRTX A6000 GPUs. The model is provided as a fully merged Hugging Face repository, meaning no additional adapter or merge steps are required for deployment.

Key Characteristics

  • Directly Loadable: The repository contains four merged model shards, including expert-SFT initialization, OPSD main-model LoRA updates, MTP LoRA updates, and directly trained full-MTP tensors.
  • Training Provenance: Detailed merge_manifest.json and training_config.json files are included, providing hashes, source identity, configuration, and finite metrics from this specific update.
  • Teacher-Only PI: During its training, the PI (Proprietary Information) was teacher-only, suggesting the model's student component should be evaluated independently without adding PI.

Usage Considerations

This model is suitable for users who require a specific snapshot of a Qwen3.5-based model after a defined training regimen. It is particularly useful for research or applications where the exact training state and integrity of the model's components are critical. Users should evaluate the model on held-out tasks, distinct from the 1024 training rows, to assess its generalization capabilities.