bmle/Qwen3-4B-Instruct_NSFW-V2.1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

bmle/Qwen3-4B-Instruct_NSFW-V2.1 is a 4 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen3-4B-Instruct-2507. This model is designed for direct inference, suitable for applications requiring a compact yet capable language model with a 32768 token context length. It is a merged full model, eliminating the need for separate LoRA adapters.

Loading preview...

Model Overview

bmle/Qwen3-4B-Instruct_NSFW-V2.1 is a 4 billion parameter instruction-tuned language model, derived from a fine-tuning of the Qwen/Qwen3-4B-Instruct-2507 base model. This version is a fully merged model, meaning it is ready for direct inference without the need to load separate LoRA adapters, simplifying deployment.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-4B-Instruct-2507.
  • Parameter Count: 4 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended responses.
  • Deployment: Provided as a merged full model, streamlining the inference process by removing the requirement for external LoRA adapters.

Use Cases

This model is suitable for developers seeking a compact, instruction-following language model for various applications where a 4B parameter size and a large context window are beneficial. Its merged nature makes it straightforward to integrate into existing workflows for direct inference.