lhpku20010120/Omni-Edu-4B
Omni-Edu-4B by lhpku20010120 is a 4.5 billion parameter language model fine-tuned from Qwen/Qwen3.5-4B-Base. It is specifically adapted using the Omni-Edu-70K dataset, suggesting a focus on educational applications or content. With a context length of 32768 tokens, it is designed for processing longer sequences of text relevant to its specialized training.
Loading preview...
Omni-Edu-4B Model Overview
Omni-Edu-4B is a 4.5 billion parameter language model developed by lhpku20010120. It is a fine-tuned variant of the Qwen/Qwen3.5-4B-Base architecture, specifically adapted for educational contexts. The model leverages a substantial 32768-token context window, enabling it to handle extensive textual inputs and maintain coherence over longer documents or conversations.
Key Characteristics
- Base Model: Fine-tuned from Qwen/Qwen3.5-4B-Base.
- Parameter Count: 4.5 billion parameters.
- Context Length: Supports a 32768-token context window.
- Training Data: Fine-tuned on the Omni-Edu-70K dataset, indicating a specialization in educational content and tasks.
Training Details
The model was trained with a learning rate of 5e-06, a total batch size of 64 (across 8 GPUs with gradient accumulation), and utilized a cosine learning rate scheduler over 3 epochs. The training environment included Transformers 5.2.0 and Pytorch 2.10.0.
Potential Use Cases
Given its fine-tuning on an educational dataset, Omni-Edu-4B is likely well-suited for applications such as:
- Generating educational content or summaries.
- Assisting with question answering in academic domains.
- Processing and understanding educational texts.
- Developing intelligent tutoring systems.