3tic/Orion-Qwen3-8B-CPT-v2607
The 3tic/Orion-Qwen3-8B-CPT-v2607 is an 8 billion parameter base model built upon Qwen3-8B-Base, specifically pre-trained (CPT) on over 40 billion tokens of Chinese and Japanese light novel data. This model is optimized for tasks requiring deep understanding and generation of content in the style and context of light novels, including translated and original works. Its specialized training makes it highly suitable for subsequent fine-tuning in creative writing, translation, and content generation within the light novel domain, supporting a 32768 token context length.
Loading preview...
Model Overview
3tic/Orion-Qwen3-8B-CPT-v2607 is an 8 billion parameter base model derived from Qwen/Qwen3-8B-Base. Its core differentiator lies in its extensive Continued Pre-Training (CPT) on a massive dataset exceeding 40 billion tokens, specifically curated from Chinese and Japanese light novel content. This specialized training regimen aims to imbue the model with a nuanced understanding of the stylistic, narrative, and linguistic characteristics prevalent in light novels.
Key Training Data Components
The model's unique capabilities stem from its diverse training data, which includes:
- Japanese Sources: Published light novel texts, web-based works, Galgame scripts, and anime subtitle texts.
- Chinese Sources: Translated light novel texts, online web novels, light novel forum translations, localized Galgame scripts, and translated anime subtitle texts. The dataset also incorporates fanfiction and other specific online content.
Use Cases and Differentiators
This model is designed as a base model for further fine-tuning, making it particularly well-suited for:
- Creative Writing: Generating narratives, dialogues, or descriptions in the style of Chinese and Japanese light novels.
- Translation: Assisting with the translation of light novel content, leveraging its domain-specific linguistic knowledge.
- Content Generation: Creating character backstories, plot outlines, or supplementary materials for light novel-related projects.
Its deep immersion in light novel data distinguishes it from general-purpose LLMs, offering enhanced performance for tasks within this specific cultural and literary domain. The model supports a substantial context length of 32768 tokens, allowing for the processing of longer narrative segments.