tanliboy/zephyr-qwen2-7b-sft

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 10, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The tanliboy/zephyr-qwen2-7b-sft model is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2-7B. It was specifically trained on the HuggingFaceH4/ultrachat_200k dataset, achieving a validation loss of 1.0646. This model is optimized for conversational AI and instruction-following tasks, leveraging its fine-tuning on a diverse chat dataset.

Loading preview...

Model Overview

tanliboy/zephyr-qwen2-7b-sft is a 7.6 billion parameter language model derived from the Qwen2-7B architecture. It has undergone supervised fine-tuning (SFT) using the comprehensive HuggingFaceH4/ultrachat_200k dataset, which is designed to enhance conversational capabilities and instruction-following.

Training Details

The model was trained with a learning rate of 2e-05 over a single epoch, utilizing a total batch size of 128 across 8 GPUs. The training process employed an Adam optimizer with standard betas and epsilon, alongside a cosine learning rate scheduler with a 0.1 warmup ratio. During training, a validation loss of 1.0646 was recorded.

Key Characteristics

  • Base Model: Qwen/Qwen2-7B
  • Fine-tuning Dataset: HuggingFaceH4/ultrachat_200k
  • Parameter Count: 7.6 billion
  • Context Length: 32768 tokens

Potential Use Cases

This model is particularly well-suited for applications requiring robust conversational abilities and adherence to user instructions, given its fine-tuning on a large-scale chat dataset. It can be applied to chatbots, virtual assistants, and other interactive AI systems.