kevincity2026/kevin_knowledge-0707-022746

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026Architecture:Transformer Featherless Exclusive Cold

kevincity2026/kevin_knowledge-0707-022746 is a 1.5 billion parameter language model merged from Qwen2.5-1.5B and Qwen2.5-Coder-1.5B-Instruct using the SLERP method. This model combines the general language understanding of Qwen2.5 with the coding capabilities of its Coder variant, making it suitable for tasks requiring both general instruction following and code-related generation or comprehension. It features a 32768 token context length, offering extended capacity for complex prompts and code snippets.

Loading preview...

Model Overview

kevincity2026/kevin_knowledge-0707-022746 is a 1.5 billion parameter language model created by kevincity2026. It was developed by merging two base models from the Qwen family: Qwen/Qwen2.5-1.5B and Qwen/Qwen2.5-Coder-1.5B-Instruct.

Merge Details

This model was constructed using the SLERP (Spherical Linear Interpolation) merge method, a technique often employed to combine the strengths of different pre-trained models. The merge configuration used a t parameter of 0.5, indicating an equal weighting between the two source models. The merge specifically targeted layers 0 through 28 of both Qwen2.5-1.5B and Qwen2.5-Coder-1.5B-Instruct.

Key Capabilities

  • Blended Expertise: Inherits general language understanding from Qwen2.5-1.5B and specialized coding abilities from Qwen2.5-Coder-1.5B-Instruct.
  • Instruction Following: Benefits from the instruction-tuned nature of the Coder variant, enhancing its ability to respond to specific prompts.
  • Extended Context: Supports a context length of 32768 tokens, allowing for processing longer inputs, including code blocks or detailed instructions.

Good For

  • Applications requiring a balance of general text generation and code-related tasks.
  • Scenarios where a smaller, efficient model with combined language and coding proficiency is needed.
  • Experiments with merged models to leverage specific strengths of different base architectures.