pltops/qwen3_5-v1
The pltops/qwen3_5-v1 model is a 9 billion parameter language model, fine-tuned from Qwen/Qwen3.5-9B. It has been specialized through training on diverse datasets including banmime, banglaprotha, scienceqa, docvqa, and mtvqa. This model is designed for tasks requiring understanding and generation based on these specific data domains, offering enhanced performance in areas covered by its training data. Its 32K context length supports processing longer inputs for complex queries.
Loading preview...
Overview
pltops/qwen3_5-v1 is a 9 billion parameter language model, fine-tuned from the base Qwen/Qwen3.5-9B architecture. This model has undergone specialized training on a combination of diverse datasets, including banmime, banglaprotha, scienceqa, docvqa, and mtvqa, to enhance its performance in specific domains.
Key Characteristics
- Base Model: Qwen/Qwen3.5-9B
- Parameter Count: 9 billion
- Context Length: 32,768 tokens
- Fine-tuning Datasets: banmime, banglaprotha, scienceqa, docvqa, mtvqa
Training Details
The model was trained with a learning rate of 0.0002, using an AdamW optimizer with a cosine learning rate scheduler. Training involved a batch size of 1 with 16 gradient accumulation steps, totaling an effective batch size of 16 over 1 epoch. The training utilized Transformers 5.14.0.dev0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.
Intended Use Cases
This model is particularly suited for applications that benefit from its specialized training on the aforementioned datasets. Developers can leverage its fine-tuned capabilities for tasks related to the content and structure found within banmime, banglaprotha, scienceqa, docvqa, and mtvqa, potentially offering improved accuracy and relevance in these areas compared to a general-purpose model.