KeefeBuild/Keefe-Discere

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 3, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Keefe-Discere is an independently developed 7.6 billion parameter instruction-following language model by KeefeBuild, built on the Qwen2.5-7B architecture with a 32K context length. It is specifically enhanced for reasoning, mathematics, programming, and agentic tool use through a two-stage process involving model merging and targeted fine-tuning. This model excels at structured problem-solving, code execution, and general instruction following, making it suitable for local, private inference.

Loading preview...

Keefe-Discere: An Enhanced Instruction-Following Model

Keefe-Discere is an independently developed 7.6 billion parameter language model by KeefeBuild, designed for robust instruction following with a strong emphasis on analytical and computational tasks. Built upon the Qwen2.5-7B architecture, it features a 32K token context window, enabling it to handle complex and lengthy interactions.

Key Capabilities

  • Advanced Reasoning: Optimized for structured problem-solving and logical deduction.
  • Mathematical Proficiency: Excels in quantitative tasks and calculations.
  • Programming & Code Execution: Highly capable in generating, debugging, and executing code, particularly Python.
  • Agentic Tool Use: Enhanced for function calling and integration with external tools.
  • General Instruction Following: Provides reliable responses across a broad range of prompts.

What Makes It Different?

Keefe-Discere distinguishes itself through a unique two-stage enhancement process. It begins with model merging (DARE-TIES), combining specialized models for general knowledge, coding, and mathematics into a balanced checkpoint. This is followed by targeted post-training (QLoRA v1.1), fine-tuning a LoRA adapter on curated data focused on instruction, reasoning, and code execution to significantly improve its agentic capabilities and mathematical accuracy. This approach results in a general-purpose, locally deployable model with a strong focus on practical, analytical applications.