philipperen55/Qwen3-14B-datasetCPT70-axolotl
The philipperen55/Qwen3-14B-datasetCPT70-axolotl is a 14 billion parameter Qwen3-Base model fine-tuned by philipperen55 using Axolotl. This model was specifically trained on the philipperen55/datasetCPT70axolotlRandomized dataset, achieving a validation loss of 1.8920 and a perplexity of 6.6325. It is optimized for tasks related to the specific data distribution of its training dataset, making it suitable for applications requiring specialized knowledge from that domain.
Loading preview...
Model Overview
This model, philipperen55/Qwen3-14B-datasetCPT70-axolotl, is a 14 billion parameter language model based on the Qwen3-Base architecture. It has been fine-tuned by philipperen55 using the Axolotl framework, leveraging Flash Attention 2 for efficient training. The model was trained for one epoch on the philipperen55/datasetCPT70axolotlRandomized dataset, with a sequence length of 2048 tokens.
Key Training Details
- Base Model: Qwen/Qwen3-14B-Base
- Training Dataset:
philipperen55/datasetCPT70axolotlRandomized - Optimizer: AdamW 8-bit with a learning rate of 4e-05
- Perplexity (PPL): Achieved 6.6325 on the evaluation set
- Validation Loss: 1.8920
- Memory Usage: Max active/allocated memory during evaluation was 137.59 GiB.
Intended Uses & Limitations
This model is specifically adapted to the data it was fine-tuned on. While the README indicates that more information is needed regarding its specific intended uses and limitations, its training on a specialized dataset suggests it would perform best on tasks aligned with that data's characteristics. Users should evaluate its performance carefully for their specific applications, especially for domains outside of its training distribution.