Jordansky/test-957-3

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Jordansky/test-957-3 is a 0.49 billion parameter causal language model from the Qwen2.5 series, developed by the Qwen Team. This base model features a 32,768 token context length and transformers architecture with RoPE, SwiGLU, and RMSNorm. It offers significantly improved capabilities in coding, mathematics, instruction following, and generating long texts, with multilingual support for over 29 languages. It is intended for further post-training applications like SFT or RLHF rather than direct conversational use.

Loading preview...

Overview

Jordansky/test-957-3 is a 0.49 billion parameter base model from the Qwen2.5 series, developed by the Qwen Team. This model is part of a new generation of Qwen large language models, ranging from 0.5B to 72B parameters, designed to offer significant improvements over Qwen2. It features a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings, supporting a full context length of 32,768 tokens.

Key Capabilities

  • Enhanced Knowledge & Skills: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Offers substantial improvements in adhering to instructions and generating long texts (over 8K tokens).
  • Structured Data Handling: Better at understanding structured data like tables and generating structured outputs, especially JSON.
  • Robustness: More resilient to diverse system prompts, enhancing role-play and chatbot condition-setting.
  • Long-Context Support: Supports up to 128K tokens and can generate up to 8K tokens.
  • Multilingual Support: Provides support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic.

Use Cases

This model is a base language model and is not recommended for direct conversational use. Instead, it is designed as a foundation for further post-training applications such as:

  • Supervised Fine-Tuning (SFT)
  • Reinforcement Learning from Human Feedback (RLHF)
  • Continued Pretraining

For more details, refer to the Qwen2.5 blog and GitHub repository.