pikaju123/Qwen2.5-3B-Instruct-Merged16bit

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The pikaju123/Qwen2.5-3B-Instruct-Merged16bit is a 7.6 billion parameter instruction-tuned causal language model, developed by pikaju123. This model is finetuned from unsloth/qwen2.5-coder-7b-instruct-bnb-4bit and was trained using Unsloth and Huggingface's TRL library for accelerated performance. It is designed for general instruction-following tasks, leveraging its Qwen2.5 architecture and a 32768 token context length.

Loading preview...

Model Overview

The pikaju123/Qwen2.5-3B-Instruct-Merged16bit is a 7.6 billion parameter instruction-tuned model developed by pikaju123. It is based on the Qwen2.5 architecture and was finetuned from the unsloth/qwen2.5-coder-7b-instruct-bnb-4bit model. A notable aspect of its development is the use of Unsloth and Huggingface's TRL library, which enabled a 2x faster training process.

Key Characteristics

  • Architecture: Qwen2.5-based, known for its strong performance across various benchmarks.
  • Parameter Count: 7.6 billion parameters, offering a balance between capability and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, beneficial for processing longer inputs and maintaining conversational coherence.
  • Training Efficiency: Leverages Unsloth for accelerated training, indicating potential for efficient deployment and fine-tuning.

Intended Use Cases

This model is suitable for a range of instruction-following applications, including:

  • General-purpose instruction following: Responding to diverse prompts and commands.
  • Text generation: Creating coherent and contextually relevant text based on instructions.
  • Conversational AI: Engaging in extended dialogues due to its large context window.

Its foundation on a coder-specific base model suggests potential strengths in code-related tasks, although the primary instruction-tuning broadens its applicability.