flytech/devchat-llama-7b
The flytech/devchat-llama-7b model is a 7 billion parameter language model fine-tuned from openlm-research/open_llama_7b_v2. It was trained on over 5000 examples across datasets like Open-Platypus, CodeAlpaca-20k, and evol-codealpaca-v1 for 3 epochs. This model is optimized for code-related tasks and general instruction following, leveraging its Llama-based architecture and 4096-token context length.
Loading preview...
Model Overview
flytech/devchat-llama-7b is a 7 billion parameter language model, fine-tuned from the openlm-research/open_llama_7b_v2 base model. This fine-tuning process involved training on over 5000 examples across three distinct datasets for 3 epochs.
Key Capabilities
- Instruction Following: Enhanced through training on diverse instruction-based datasets.
- Code Generation: Benefits from inclusion of code-centric datasets like
sahil2801/CodeAlpaca-20kandtheblackcat102/evol-codealpaca-v1. - General Language Understanding: Built upon the robust Open Llama 7B v2 architecture.
Training Details
The model was trained with a learning rate of 0.0002, a batch size of 8 (total effective batch size of 16 with gradient accumulation), and utilized an Adam optimizer. The training spanned 3 epochs, focusing on improving performance across a range of tasks, particularly those involving code and detailed instructions. The context length for this model is 4096 tokens.
Intended Use Cases
This model is suitable for applications requiring:
- Responding to complex instructions.
- Assisting with code generation and understanding.
- General natural language processing tasks where a 7B parameter model is appropriate.