MINZIK77/lm-sft-tulu-3b-ckpts
MINZIK77/lm-sft-tulu-3b-ckpts is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B-Instruct with a context length of 32768 tokens. This model was trained using Supervised Fine-Tuning (SFT) via the TRL framework. It is designed for general text generation tasks, leveraging the capabilities of its base Qwen2.5-3B-Instruct architecture.
Loading preview...
Model Overview
MINZIK77/lm-sft-tulu-3b-ckpts is a 3.1 billion parameter language model, fine-tuned from the robust Qwen/Qwen2.5-3B-Instruct base model. It features a substantial context length of 32768 tokens, making it suitable for processing longer inputs and generating extended responses. The model was developed using Supervised Fine-Tuning (SFT) techniques, leveraging the TRL (Transformers Reinforcement Learning) framework.
Key Capabilities
- Instruction Following: Inherits and refines the instruction-following capabilities of its Qwen2.5-3B-Instruct base.
- Text Generation: Capable of generating coherent and contextually relevant text based on user prompts.
- Extended Context: Benefits from a 32K token context window, allowing for more detailed and comprehensive interactions.
Training Details
The model underwent a Supervised Fine-Tuning (SFT) process. The training utilized specific versions of popular machine learning frameworks:
- TRL: 1.7.1
- Transformers: 4.57.6
- Pytorch: 2.10.0
- Datasets: 4.7.0
- Tokenizers: 0.22.2
Good For
- General-purpose text generation tasks.
- Applications requiring a model with a relatively small parameter count but a large context window.
- Developers looking for a fine-tuned Qwen2.5-3B-Instruct variant for specific SFT-driven use cases.