Toleng/koplak-flash-3b-instruct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Toleng/koplak-flash-3b-instruct is a 3.1 billion parameter instruction-tuned causal language model developed by Toleng. Finetuned from unsloth/qwen2.5-coder-3b-instruct-bnb-4bit, this model leverages Unsloth and Huggingface's TRL library for accelerated training. It is optimized for instruction-following tasks, making it suitable for applications requiring responsive and accurate text generation. The model supports a context length of 32768 tokens, enhancing its ability to handle longer prompts and maintain conversational coherence.

Loading preview...

Toleng/koplak-flash-3b-instruct Overview

Toleng/koplak-flash-3b-instruct is a 3.1 billion parameter instruction-tuned language model developed by Toleng. It is finetuned from the unsloth/qwen2.5-coder-3b-instruct-bnb-4bit base model, indicating a potential specialization or strong performance in coding-related instruction following, though the README does not explicitly detail its specific coding capabilities.

Key Characteristics

  • Architecture: Based on the Qwen2.5 model family.
  • Parameter Count: 3.1 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing for processing and generating longer sequences of text.
  • Training Efficiency: The model was trained using Unsloth and Huggingface's TRL library, which enabled 2x faster finetuning.
  • License: Distributed under the Apache-2.0 license, providing broad usage permissions.

Good For

  • Instruction Following: Designed for tasks that require the model to adhere to specific instructions.
  • Applications requiring efficient inference: Its 3.1B parameter size makes it suitable for scenarios where faster processing is crucial.
  • Developers leveraging Unsloth: Demonstrates the effectiveness of Unsloth for accelerated model finetuning.