ArchiveStudio/Qwen2.5-7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ArchiveStudio/Qwen2.5-7B is a 7.61 billion parameter causal language model from the Qwen2.5 series, developed by Qwen Team. This base model features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a substantial context length of 131,072 tokens. It offers significantly improved knowledge, coding, and mathematical capabilities compared to its predecessors, alongside enhanced instruction following and structured data understanding. It is designed for further fine-tuning, such as SFT or RLHF, rather than direct conversational use.

Loading preview...

Overview

ArchiveStudio/Qwen2.5-7B is a 7.61 billion parameter base causal language model, part of the Qwen2.5 series developed by the Qwen Team. This model builds upon the Qwen2 architecture, incorporating improvements in knowledge, coding, and mathematics through specialized expert models. It features a transformer architecture with RoPE, SwiGLU, RMSNorm, and Attention QKV bias, and supports an extensive context length of 131,072 tokens.

Key Capabilities

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in general knowledge, coding, and mathematics.
  • Instruction Following: Demonstrates substantial improvements in adhering to instructions and generating structured outputs like JSON.
  • Long-Context Support: Capable of processing inputs up to 128K tokens and generating outputs up to 8K tokens.
  • Multilingual Support: Offers support for over 29 languages, including major global languages.
  • Structured Data Understanding: Better at interpreting structured data, such as tables.

Good For

  • Foundation for Fine-tuning: Ideal for developers looking to apply post-training methods like Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining.
  • Applications Requiring Long Context: Suitable for tasks that benefit from processing very long input sequences.
  • Developing Specialized Models: Its enhanced base capabilities in coding and mathematics make it a strong starting point for domain-specific applications in these areas.