cuong1692001/Terminal_complete_12k
Terminal-complete_12k is an 8 billion parameter language model fine-tuned by cuong1692001, based on the Qwen3-8B architecture. It was trained on the qwen_data_complete dataset, suggesting a focus on comprehensive language understanding and generation. This model is optimized for general-purpose text completion tasks, leveraging its 32768-token context window for handling extensive inputs.
Loading preview...
Model Overview
Terminal-complete_12k is an 8 billion parameter language model developed by cuong1692001. It is a fine-tuned variant of the Qwen3-8B base model, specifically trained on the qwen_data_complete dataset. This fine-tuning process aims to enhance its capabilities for a broad range of text completion and generation tasks.
Training Details
The model underwent training with the following key hyperparameters:
- Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH with default betas and epsilon
- Scheduler: Cosine learning rate scheduler
- Epochs: 2.0
- Batch Size: A total training batch size of 4 was used across 4 GPUs.
Framework Versions
The training environment utilized:
- Transformers 5.6.0
- Pytorch 2.11.0+cu130
- Datasets 4.0.0
- Tokenizers 0.22.2
Potential Use Cases
Given its foundation on Qwen3-8B and training on a 'complete' dataset, Terminal-complete_12k is likely suitable for:
- General text generation and completion
- Assisting with various language-based tasks requiring broad knowledge
- Applications benefiting from a 32k context window for longer inputs.