aldandev/viper-spark

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

aldandev/viper-spark is a 1.54 billion parameter base causal language model from the Qwen2.5 series, developed by the Qwen Team. This model features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a context length of 32,768 tokens. It offers significant improvements in coding, mathematics, instruction following, and structured data understanding compared to its predecessors. As a base model, it is designed for further fine-tuning, such as SFT or RLHF, rather than direct conversational use.

Loading preview...

Qwen2.5-1.5B: A Foundation for Advanced LLM Applications

aldandev/viper-spark is the 1.54 billion parameter base model from the Qwen2.5 series, developed by the Qwen Team. This model builds upon the Qwen2 architecture, incorporating significant enhancements across several key areas. It is designed as a pre-trained causal language model, making it suitable for developers looking to build specialized applications through further fine-tuning.

Key Capabilities and Improvements

  • Enhanced Knowledge and Reasoning: Features significantly improved capabilities in coding and mathematics, benefiting from specialized expert models.
  • Instruction Following: Demonstrates substantial improvements in adhering to instructions, generating long texts (over 8K tokens), and understanding structured data like tables.
  • Structured Output Generation: Excels at producing structured outputs, particularly JSON, and is more resilient to diverse system prompts, aiding in robust chatbot development and role-play scenarios.
  • Long-Context Support: While the base model supports a context length of 32,768 tokens, the Qwen2.5 series generally supports up to 128K tokens and can generate up to 8K tokens.
  • Multilingual Support: Offers broad multilingual capabilities, supporting over 29 languages including Chinese, English, French, Spanish, German, Japanese, and Korean.

Model Architecture and Specifications

This specific model is a causal language model with 1.54 billion parameters (1.31 billion non-embedding parameters), 28 layers, and 12 attention heads (with 2 for KV in GQA). It utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.

Recommended Use

As a base language model, aldandev/viper-spark is not recommended for direct conversational use. Instead, it serves as an excellent foundation for post-training techniques such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining to adapt it for specific downstream tasks and conversational agents.