cuong1692001/Terminal-16k-bottom50
Terminal-16k-bottom50 is an 8 billion parameter language model developed by cuong1692001, fine-tuned from Terminal-complete_8k. This model is specifically trained on the nemotron_complete_bottom_50_16k dataset, suggesting an optimization for tasks related to its training data. It supports a context length of 32768 tokens, making it suitable for processing extensive inputs.
Loading preview...
Overview
Terminal-16k-bottom50 is an 8 billion parameter language model, fine-tuned by cuong1692001. It is based on the Terminal-complete_8k architecture and has been specialized through training on the nemotron_complete_bottom_50_16k dataset. This model is designed to handle a substantial context window of 32768 tokens, allowing for the processing of longer texts and more complex queries.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 1e-05
- Batch Size: 1 (train), 8 (eval)
- Optimizer: ADAMW_TORCH with default betas and epsilon
- LR Scheduler: Cosine
- Epochs: 2.0
Intended Uses & Limitations
Specific intended uses and limitations are not detailed in the provided information. Users should evaluate its performance on tasks relevant to the nemotron_complete_bottom_50_16k dataset to determine suitability. The model's large context window suggests potential for applications requiring extensive input understanding or generation.