ChesterProgrammer/CPT-Lucy-V0
CPT-Lucy-V0 is a 9 billion parameter Qwen3.5-based language model developed by ChesterProgrammer, continued pre-trained for 3 epochs. This model was finetuned from DreamFast/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Safetensor-Benchmark, utilizing Unsloth and Huggingface's TRL library for accelerated training. It features a 32768 token context length and was trained with specific LoRA configurations including a rank of 32 and alpha of 64.
Loading preview...
CPT-Lucy-V0: A Continued Pre-trained Qwen3.5 Model
CPT-Lucy-V0 is a 9 billion parameter language model developed by ChesterProgrammer, built upon the Qwen3.5 architecture. This model underwent continued pre-training for 3 epochs, leveraging the DreamFast/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Safetensor-Benchmark as its base. A notable aspect of its development is the use of Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
Key Training Details
The model's training regimen included:
- Method: Continued pretraining over 3 epochs.
- Optimizer: AdamW 8-bit with a learning rate of 0.00005.
- Context Length: Trained with a context length of 2048 tokens, though the model supports 32768 tokens.
- LoRA Configuration: Utilized LoRA with a rank of 32, alpha of 64, and 0.0 dropout.
- Dataset: Trained on a 2-Pass CPT Corpus.
What Makes This Different?
CPT-Lucy-V0 distinguishes itself through its optimized training methodology, achieving faster iteration cycles thanks to Unsloth. Its continued pre-training from an "Uncensored-HauhauCS-Aggressive" base suggests a potential focus on broader or less restricted content generation compared to standard models. The specific LoRA parameters indicate a fine-tuned approach to adapting the base model's capabilities.
Should I Use This?
This model is suitable for users looking for a Qwen3.5-based model that has undergone additional pre-training with specific LoRA configurations. Its origins from an "Uncensored-HauhauCS-Aggressive" base might appeal to use cases requiring less constrained output. Developers interested in models optimized for faster training with Unsloth may also find this a relevant option.