ToastyPigeon/gemma-4-12b-full-cpt-itvec
ToastyPigeon/gemma-4-12b-full-cpt-itvec is a 12 billion parameter Gemma-4 derivative model with a 32768 token context length. It combines a prose/story-heavy CPT merge with an instruct task vector to retain instruction-following while enhancing prose generation. This model is specifically designed for text-completion, storywriting experiments, and as a base for downstream roleplay or style-focused finetunes.
Loading preview...
Model Overview
ToastyPigeon/gemma-4-12b-full-cpt-itvec is a 12 billion parameter model derived from the google/gemma-4-12b base. It integrates a full-CPT (Continual Pre-Training) merge with an instruct task vector to achieve a unique blend of capabilities. The CPT step involved training on a diverse mix of prose and chat-log corpora, including datasets focused on story and erotic fiction, to enhance its prose generation and stylistic output.
Key Capabilities
- Enhanced Prose Generation: Benefits from a CPT merge on story-heavy and prose datasets, making it adept at generating narrative and stylistic text.
- Instruction Following: Incorporates an instruct task vector to maintain strong instruction-following and formatting behavior, which is often lost in pure CPT models.
- Unified Multimodal Architecture: Retains the
Gemma4UnifiedForConditionalGenerationarchitecture, with vision/audio towers unchanged from the base model.
Intended Use Cases
- Text Completion and Storywriting: Ideal for experimental text generation, particularly in creative writing and story development.
- Base for Finetuning: Serves as an excellent starting point for further finetuning, especially for roleplay (RP) or specific style adaptations that require both strong prose and instruct-style behavior.
Content Note: Parts of the CPT training data include adult/erotic fiction. This model is intended for fiction-writing research and style experimentation by adults and is not suitable for production deployment or use by minors.