Toleng/koplak-flash-1.5b-instruct
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
Toleng/koplak-flash-1.5b-instruct is a 1.5 billion parameter instruction-tuned Qwen2 model developed by Toleng, fine-tuned from unsloth/Qwen2.5-Coder-1.5B-Instruct-bnb-4bit. This model was trained significantly faster using Unsloth and Huggingface's TRL library, offering a 32768 token context length. It is designed for efficient performance in instruction-following tasks, leveraging its optimized training process.
Loading preview...
Toleng/koplak-flash-1.5b-instruct Overview
This model is a 1.5 billion parameter instruction-tuned variant of the Qwen2 architecture, developed by Toleng. It was fine-tuned from the unsloth/Qwen2.5-Coder-1.5B-Instruct-bnb-4bit base model, indicating a potential focus on coding or instruction-following tasks.
Key Characteristics
- Architecture: Based on the Qwen2 model family.
- Parameter Count: Features 1.5 billion parameters, making it a relatively compact yet capable model.
- Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs and maintaining conversational history.
- Optimized Training: A key differentiator is its training methodology; it was trained approximately 2 times faster using the Unsloth library in conjunction with Huggingface's TRL library. This optimization suggests a focus on efficiency and rapid iteration.
Potential Use Cases
- Instruction Following: Given its instruction-tuned nature, it is well-suited for tasks requiring adherence to specific commands or prompts.
- Resource-Efficient Applications: The combination of its 1.5B parameter size and optimized training makes it a candidate for applications where computational resources or inference speed are critical.
- Experimental Development: The use of Unsloth for faster training could make this model appealing for developers looking to quickly prototype and experiment with instruction-tuned LLMs.