SHIKARI2/Malvos-32B-Merged
TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
SHIKARI2/Malvos-32B-Merged is a 32.8 billion parameter Qwen2-based causal language model developed by SHIKARI2. This model was finetuned from unsloth/DeepSeek-R1-Distill-Qwen-32B-unsloth-bnb-4bit using Unsloth and Huggingface's TRL library, enabling faster training. With a 32768 token context length, it is suitable for applications requiring extensive context processing.
Loading preview...
Overview
SHIKARI2/Malvos-32B-Merged is a 32.8 billion parameter language model based on the Qwen2 architecture. Developed by SHIKARI2, this model was finetuned from the unsloth/DeepSeek-R1-Distill-Qwen-32B-unsloth-bnb-4bit base model.
Key Characteristics
- Architecture: Qwen2-based, a causal language model.
- Parameter Count: 32.8 billion parameters.
- Context Length: Supports a substantial context window of 32768 tokens.
- Training Efficiency: Finetuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
Potential Use Cases
- Applications requiring a large context window for processing long documents or conversations.
- Tasks benefiting from a Qwen2-based model's general language understanding and generation capabilities.
- Scenarios where efficient finetuning methods, such as those provided by Unsloth, are valued for model development.