cuong1692001/Terminal-12k-bottom80
Terminal-12k-bottom80 is an 8 billion parameter language model developed by cuong1692001, fine-tuned from the Terminal-complete_8k base model. This model was trained on the nemotron_complete_bottom_80_12k dataset, suggesting a specialization in processing specific data distributions or tasks. With a context length of 32768 tokens, it is designed for applications requiring extensive contextual understanding.
Loading preview...
Model Overview
Terminal-12k-bottom80 is an 8 billion parameter language model, fine-tuned by cuong1692001. It is based on the existing /helios-storage/helios4-data/cuong/Terminal-complete_8k model and has been further trained on the nemotron_complete_bottom_80_12k dataset. This fine-tuning process indicates a potential specialization for tasks related to the characteristics of this specific dataset.
Key Training Details
The model underwent training with the following hyperparameters:
- Learning Rate: 1e-05
- Batch Sizes:
train_batch_sizeof 1,eval_batch_sizeof 8 (total effective batch sizes of 4 and 32 respectively across 4 GPUs) - Optimizer: ADAMW_TORCH with default betas and epsilon
- Scheduler: Cosine learning rate scheduler
- Epochs: 2.0
Intended Uses & Limitations
Specific intended uses and limitations are not detailed in the provided information, suggesting further evaluation or documentation is needed to fully understand its optimal applications and potential constraints. Developers should consider the model's training data and fine-tuning approach when determining suitability for their specific use cases.