SHIKARI2/Malvos-32B-Merged

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SHIKARI2/Malvos-32B-Merged is a 32.8 billion parameter Qwen2-based causal language model developed by SHIKARI2. This model was finetuned from unsloth/DeepSeek-R1-Distill-Qwen-32B-unsloth-bnb-4bit using Unsloth and Huggingface's TRL library, enabling faster training. With a 32768 token context length, it is suitable for applications requiring extensive context processing.

Loading preview...

Overview

SHIKARI2/Malvos-32B-Merged is a 32.8 billion parameter language model based on the Qwen2 architecture. Developed by SHIKARI2, this model was finetuned from the unsloth/DeepSeek-R1-Distill-Qwen-32B-unsloth-bnb-4bit base model.

Key Characteristics

  • Architecture: Qwen2-based, a causal language model.
  • Parameter Count: 32.8 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Training Efficiency: Finetuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.

Potential Use Cases

  • Applications requiring a large context window for processing long documents or conversations.
  • Tasks benefiting from a Qwen2-based model's general language understanding and generation capabilities.
  • Scenarios where efficient finetuning methods, such as those provided by Unsloth, are valued for model development.