Pranjalps1/Qwen3.5-2b-code

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 18, 2026Architecture:Transformer Featherless Exclusive Cold

Pranjalps1/Qwen3.5-2b-code is a 2.3 billion parameter Qwen3.5 model, fine-tuned and converted to GGUF format using Unsloth. This model is optimized for efficient deployment and inference, supporting a 32768 token context length. It is designed for general language tasks, with specific GGUF files available for both text-only and multimodal applications.

Loading preview...

Qwen3.5-2b-code: GGUF Model

This model, developed by Pranjalps1, is a 2.3 billion parameter variant of the Qwen3.5 architecture. It has been specifically fine-tuned and converted into the GGUF format using the Unsloth framework, which facilitates faster training and efficient deployment on various hardware.

Key Capabilities

  • Efficient Inference: Provided in GGUF format, enabling optimized performance with llama.cpp and similar tools.
  • Multimodal Support: Includes a BF16-mmproj.gguf file, indicating potential for multimodal applications alongside a text-only variant.
  • Unsloth Optimization: Benefits from Unsloth's training optimizations, allowing for faster fine-tuning processes.

Good For

  • Local Deployment: Ideal for running on consumer hardware due to its GGUF format and relatively compact size.
  • Experimentation: Suitable for developers looking to experiment with Qwen3.5 models in a GGUF environment.
  • Text and Multimodal Tasks: Offers flexibility for both pure text generation and tasks requiring multimodal input processing, depending on the chosen GGUF file.