KIng241300/promptforge-nano-7b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

KIng241300/promptforge-nano-7b is a 7.61 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This model significantly enhances capabilities in coding, mathematics, and instruction following, while also improving long text generation up to 8K tokens and structured data understanding. It supports a full context length of 131,072 tokens and is multilingual, covering over 29 languages.

Loading preview...

Qwen2.5-7B-Instruct: An Enhanced 7.6 Billion Parameter Model

This model, KIng241300/promptforge-nano-7b, is an instruction-tuned variant from the Qwen2.5 series, building upon the Qwen2 architecture. Developed by Qwen, it features 7.61 billion parameters and is designed as a causal language model.

Key Capabilities and Improvements

Qwen2.5 introduces several significant advancements over its predecessor:

  • Enhanced Knowledge & Reasoning: Greatly improved performance in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Demonstrates significant improvements in adhering to instructions and generating structured outputs, including JSON.
  • Long Context & Generation: Supports an extensive context length of up to 131,072 tokens and can generate texts up to 8,192 tokens. It utilizes YaRN for handling long texts, with a default config.json set for 32,768 tokens.
  • Multilingual Support: Offers robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and Japanese.
  • Structured Data Understanding: Better at interpreting structured data such as tables and more resilient to diverse system prompts for role-play scenarios.

Architecture and Features

The model employs a transformer architecture with RoPE, SwiGLU, RMSNorm, and Attention QKV bias. It has 28 layers and 28 attention heads (with 4 for KV in GQA).

When to Use This Model

This model is particularly well-suited for applications requiring:

  • Complex Code Generation and Mathematical Problem Solving.
  • Precise Instruction Following and generating specific output formats.
  • Processing and Generating Long Documents or conversations.
  • Multilingual Applications across a broad range of languages.
  • Chatbot implementations that require robust role-play and condition-setting capabilities.