artificialguybr/QWEN-2.5-0.5B-Synthia-II

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 28, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

artificialguybr/QWEN-2.5-0.5B-Synthia-II is a 490 million parameter language model, fine-tuned by artificialguybr from the Qwen2.5-0.5B base model. Optimized for conversational AI and instruction following, it leverages a 32K token context window and advanced architectural features like RoPE and SwiGLU. This model excels at generating coherent text and handling multi-turn dialogues, making it suitable for interactive AI applications.

Loading preview...

Model Overview

artificialguybr/QWEN-2.5-0.5B-Synthia-II is a fine-tuned version of the Qwen2.5-0.5B base model, developed by artificialguybr. This model features 490 million parameters (360M non-embedding) and is built with 24 transformer layers, 14 query attention heads, and 2 key/value heads (GQA architecture). It supports a 32,768 token context length and incorporates advanced features such as RoPE positional embeddings, SwiGLU activations, and RMSNorm.

Key Differentiator

The primary distinction of this model is its fine-tuning on the Synthia-v1.5-II dataset, specifically designed to enhance instruction following and conversational abilities. The training process involved careful hyperparameter tuning over 3 epochs, using a learning rate of 1e-05 and a sequence length of 4096, to optimize for natural dialogue and instruction adherence.

Intended Uses

  • Conversational AI applications
  • Instruction following tasks
  • Text generation with strong coherence
  • Multi-turn dialogue systems

Limitations

As a 0.5B parameter model, it may not perform as well as larger models on complex reasoning tasks. Performance in non-English languages may also be limited, and users should be aware of potential biases inherited from the training data.