xiaolesu/OsmosisProofling-SFT-NT-GRPO-TK
The xiaolesu/OsmosisProofling-SFT-NT-GRPO-TK is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen3-8B. Developed by xiaolesu, it leverages the OsmosisProofling-v3-SFT dataset and Axolotl for training. This model is optimized for general language understanding and generation tasks, demonstrating a validation perplexity of 1.4252.
Loading preview...
Model Overview
The xiaolesu/OsmosisProofling-SFT-NT-GRPO-TK is an 8 billion parameter language model, fine-tuned from the Qwen/Qwen3-8B base model. It was trained using the Axolotl framework, incorporating specific Liger plugin enhancements for Rope, RMS Norm, GLU activation, and Layer Norm, suggesting a focus on efficient and potentially improved architectural components.
Key Characteristics
- Base Model: Qwen/Qwen3-8B, a robust foundation for general language tasks.
- Fine-tuning Dataset: Utilizes
xiaolesu/OsmosisProofling-v3-SFTfor instruction-following capabilities. - Training Configuration: Employs a sequence length of 4096 tokens with sample packing and flex attention, indicating an optimization for handling longer contexts and efficient processing.
- Performance: Achieved a validation loss of 0.3543 and a perplexity (Ppl) of 1.4252 on its evaluation set, demonstrating strong performance in language modeling.
Potential Use Cases
This model is suitable for a variety of natural language processing applications where a Qwen3-8B-based instruction-tuned model is beneficial. Its fine-tuning on the OsmosisProofling-v3-SFT dataset suggests proficiency in tasks aligned with that dataset's content, likely involving instruction-following and general conversational abilities. The architectural enhancements from the Liger plugin may contribute to its efficiency and stability.