MaziyarPanahi/calme-2.1-phi3-4b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:4kPublished:May 9, 2024License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

MaziyarPanahi/calme-2.1-phi3-4b is a 4 billion parameter language model fine-tuned by MaziyarPanahi using DPO on Microsoft's Phi-3-mini-4k-instruct architecture. This model is designed for instruction-following tasks, leveraging its base model's capabilities. It utilizes a 4096-token context length and is suitable for general conversational AI applications requiring a compact yet capable model.

Loading preview...

Model Overview

MaziyarPanahi/calme-2.1-phi3-4b is a 4 billion parameter language model developed by MaziyarPanahi. It is a fine-tuned version of microsoft/Phi-3-mini-4k-instruct, utilizing Direct Preference Optimization (DPO) for enhanced performance. This model is built upon the Phi-3 architecture, known for its efficiency and capability in smaller parameter counts.

Key Capabilities and Performance

This model is designed for instruction-following and general conversational tasks. Its performance has been evaluated on the Open LLM Leaderboard, showing a balanced average score across various benchmarks. Notable scores include:

  • IFEval (0-Shot): 55.25
  • BBH (3-Shot): 38.12
  • MMLU (5-Shot): 68.96
  • GSM8k (5-Shot): 72.25

These metrics indicate its proficiency in reasoning, common sense, and mathematical problem-solving relative to its size.

Prompt Template

The model uses the ChatML prompt format, which is standard for many instruction-tuned models, facilitating clear role separation between system, user, and assistant messages.

Use Cases

This model is suitable for applications requiring a compact yet effective language model for tasks such as:

  • General-purpose chatbots
  • Instruction-following agents
  • Text generation where efficiency and moderate performance are key