M4-ai/Hercules-Qwen1.5-14B

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0216 / Output $1.12Concurrent Unit Cost:1Model Size:14.2BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 30, 2024License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

Hercules-Qwen1.5-14B is a 14.2 billion parameter language model developed by M4-ai, fine-tuned from Qwen1.5-14B. It is optimized for a broad range of tasks including math, coding, function calling, and roleplay, utilizing 700,000 examples from the Hercules-v4 dataset. This model offers general-purpose assistant capabilities and supports a context length of 32768 tokens.

Loading preview...

Hercules-Qwen1.5-14B Overview

Hercules-Qwen1.5-14B is a 14.2 billion parameter language model developed by M4-ai, building upon the Qwen1.5-14B architecture. This model has been extensively fine-tuned using 700,000 examples from the Hercules-v4 dataset, enhancing its capabilities across various domains.

Key Capabilities

  • Mathematical Reasoning: Demonstrates proficiency in solving mathematical problems.
  • Code Generation: Capable of generating and understanding code.
  • Function Calling: Supports function calling mechanisms for integration with external tools.
  • Roleplay: Excels in conversational roleplay scenarios.
  • General Purpose Assistant: Functions effectively as a versatile assistant for diverse queries.
  • Question Answering: Provides accurate answers to a wide range of questions.
  • Chain-of-Thought: Supports complex reasoning through chain-of-thought processes.

Training Details

The model was fine-tuned using the Hercules-v4.0 dataset with a bf16 non-mixed precision training regime. The training utilized 8 Kaggle TPUs, with a global batch size of 128 and a sequence length of 1024.

Good For

This model is suitable for developers and researchers looking for a robust, general-purpose language model with strong performance in specialized areas like coding, math, and function calling. Its broad fine-tuning makes it adaptable for various assistant-like applications and complex reasoning tasks.