ayh015/myLightningOPD
The ayh015/myLightningOPD model is a 4 billion parameter language model fine-tuned from the Qwen3-4B-Base architecture. It was trained on the openthoughts3_300k_qwen3-8b dataset, suggesting a specialization in processing or generating content related to open thoughts or specific dialogue. With a context length of 32768 tokens, it is designed for applications requiring extensive contextual understanding and generation.
Loading preview...
Model Overview
The ayh015/myLightningOPD model is a 4 billion parameter language model, fine-tuned from the Qwen3-4B-Base architecture. This model was specifically trained on the openthoughts3_300k_qwen3-8b dataset, indicating a potential specialization in tasks related to processing or generating content within a specific domain, possibly involving open-ended discussions or thought processes. It supports a substantial context length of 32768 tokens, allowing it to handle long inputs and generate coherent, extended outputs.
Training Details
The model underwent 3000 training steps using the following key hyperparameters:
- Learning Rate: 8e-05
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Batch Size: A total training batch size of 32 (4 per device across 4 GPUs with 2 gradient accumulation steps)
- LR Scheduler: Cosine with a 0.1 warmup ratio
Intended Use Cases
While specific intended uses are not detailed in the provided information, its fine-tuning on the openthoughts3_300k_qwen3-8b dataset suggests potential applications in:
- Dialogue systems: Engaging in open-ended conversations.
- Content generation: Creating text that reflects diverse perspectives or thought processes.
- Contextual understanding: Analyzing and summarizing long documents or discussions due to its large context window.