SamMikaelson/Qwen3-4B-Hermes-16bit

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 9, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SamMikaelson/Qwen3-4B-Hermes-16bit is a 4 billion parameter Qwen3 model developed by SamMikaelson, fine-tuned from unsloth/qwen3-4b-instruct-2507-unsloth-bnb-4bit. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training. With a context length of 32768 tokens, it is optimized for efficient instruction-following tasks.

Loading preview...

Model Overview

SamMikaelson/Qwen3-4B-Hermes-16bit is a 4 billion parameter language model developed by SamMikaelson. It is a fine-tuned variant of the Qwen3 architecture, specifically building upon the unsloth/qwen3-4b-instruct-2507-unsloth-bnb-4bit base model. The model benefits from an extended context length of 32768 tokens, allowing it to process longer sequences of text.

Key Training Details

This model was trained with a focus on efficiency, utilizing:

  • Unsloth: A library known for accelerating the training process of large language models, enabling 2x faster training for this specific Qwen3 iteration.
  • Huggingface's TRL library: The Transformer Reinforcement Learning (TRL) library was employed, suggesting an instruction-tuned or alignment-focused training approach.

Intended Use Cases

Given its instruction-tuned nature and efficient training methodology, this model is suitable for:

  • Instruction Following: Executing commands and responding to prompts effectively.
  • Efficient Deployment: Its 4 billion parameter size, combined with optimized training, makes it a candidate for applications where computational resources are a consideration.

License

The model is released under the Apache-2.0 license.