etiennebamas/qwen3-sft-classic-small-data
etiennebamas/qwen3-sft-classic-small-data is an 8 billion parameter language model, fine-tuned from formalmathatepfl/qwen3-cpt on a specific SFT dataset. This model is designed for general language tasks, leveraging its Qwen3 base architecture and a 32K context length. Its fine-tuning process aims to enhance performance on instruction-following and conversational applications.
Loading preview...
Model Overview
This model, etiennebamas/qwen3-sft-classic-small-data, is an 8 billion parameter language model derived from formalmathatepfl/qwen3-cpt. It has undergone supervised fine-tuning (SFT) using a dedicated dataset, aiming to improve its capabilities in instruction-following and general conversational tasks. The base model's architecture, Qwen3, provides a robust foundation for diverse language generation and understanding.
Key Characteristics
- Base Model: Fine-tuned from
formalmathatepfl/qwen3-cpt. - Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
- Context Length: Supports a substantial context window of 32,768 tokens, enabling processing of longer inputs and generating coherent, extended responses.
- Training: Fine-tuned with a learning rate of 2e-05, using an AdamW optimizer and a cosine learning rate scheduler over 1 epoch.
Intended Use Cases
This model is suitable for applications requiring a capable language model with a focus on instruction adherence and generating human-like text. Its fine-tuning on an SFT dataset suggests improved performance in:
- Instruction Following: Responding accurately to given prompts and instructions.
- Conversational AI: Engaging in more natural and coherent dialogues.
- Text Generation: Creating various forms of text content based on input.
Further details on specific intended uses and limitations would require more information from the original model card.