lhpku20010120/Omni-Edu-4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 16, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

Omni-Edu-4B by lhpku20010120 is a 4.5 billion parameter language model fine-tuned from Qwen/Qwen3.5-4B-Base. It is specifically adapted using the Omni-Edu-70K dataset, suggesting a focus on educational applications or content. With a context length of 32768 tokens, it is designed for processing longer sequences of text relevant to its specialized training.

Loading preview...

Omni-Edu-4B Model Overview

Omni-Edu-4B is a 4.5 billion parameter language model developed by lhpku20010120. It is a fine-tuned variant of the Qwen/Qwen3.5-4B-Base architecture, specifically adapted for educational contexts. The model leverages a substantial 32768-token context window, enabling it to handle extensive textual inputs and maintain coherence over longer documents or conversations.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3.5-4B-Base.
  • Parameter Count: 4.5 billion parameters.
  • Context Length: Supports a 32768-token context window.
  • Training Data: Fine-tuned on the Omni-Edu-70K dataset, indicating a specialization in educational content and tasks.

Training Details

The model was trained with a learning rate of 5e-06, a total batch size of 64 (across 8 GPUs with gradient accumulation), and utilized a cosine learning rate scheduler over 3 epochs. The training environment included Transformers 5.2.0 and Pytorch 2.10.0.

Potential Use Cases

Given its fine-tuning on an educational dataset, Omni-Edu-4B is likely well-suited for applications such as:

  • Generating educational content or summaries.
  • Assisting with question answering in academic domains.
  • Processing and understanding educational texts.
  • Developing intelligent tutoring systems.