flytech/devchat-llama-7b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 5, 2023License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The flytech/devchat-llama-7b model is a 7 billion parameter language model fine-tuned from openlm-research/open_llama_7b_v2. It was trained on over 5000 examples across datasets like Open-Platypus, CodeAlpaca-20k, and evol-codealpaca-v1 for 3 epochs. This model is optimized for code-related tasks and general instruction following, leveraging its Llama-based architecture and 4096-token context length.

Loading preview...

Model Overview

flytech/devchat-llama-7b is a 7 billion parameter language model, fine-tuned from the openlm-research/open_llama_7b_v2 base model. This fine-tuning process involved training on over 5000 examples across three distinct datasets for 3 epochs.

Key Capabilities

  • Instruction Following: Enhanced through training on diverse instruction-based datasets.
  • Code Generation: Benefits from inclusion of code-centric datasets like sahil2801/CodeAlpaca-20k and theblackcat102/evol-codealpaca-v1.
  • General Language Understanding: Built upon the robust Open Llama 7B v2 architecture.

Training Details

The model was trained with a learning rate of 0.0002, a batch size of 8 (total effective batch size of 16 with gradient accumulation), and utilized an Adam optimizer. The training spanned 3 epochs, focusing on improving performance across a range of tasks, particularly those involving code and detailed instructions. The context length for this model is 4096 tokens.

Intended Use Cases

This model is suitable for applications requiring:

  • Responding to complex instructions.
  • Assisting with code generation and understanding.
  • General natural language processing tasks where a 7B parameter model is appropriate.