apolloransom/Qwen2.5-Coder-7B-Instruct-MLX

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

apolloransom/Qwen2.5-Coder-7B-Instruct-MLX is an MLX-optimized version of the 7.6 billion parameter Qwen2.5-Coder-7B-Instruct model, designed for efficient local execution on Apple Silicon. This model specializes in code generation and assistance, featuring a 32K context window. It is pre-configured for seamless integration with tools like oMLX and Zed IDE, providing privacy-focused coding support.

Loading preview...

Overview

This model, apolloransom/Qwen2.5-Coder-7B-Instruct-MLX, is an MLX-optimized variant of the 7.6 billion parameter Qwen2.5-Coder-7B-Instruct model. It is specifically engineered for high-performance local inference on Apple Silicon (macOS) devices.

Key Capabilities

  • Code Generation & Assistance: Optimized for coding tasks, leveraging the base model's capabilities.
  • MLX Optimization: Provides efficient execution on Apple Silicon hardware.
  • 32K Context Window: Supports extended conversational and code contexts.
  • Zero-Configuration Integration: Pre-bundled with chat templates and tokenizer configurations for direct use with specific development environments.
  • OpenAI-Compatible API Support: Designed to work with tools like oMLX, which exposes an OpenAI-compatible API for seamless integration with IDEs like Zed.
  • Privacy-Focused: Ideal for local execution, ensuring code and data remain on the user's device.

Good For

  • Local Coding Assistance: Developers seeking an on-device AI assistant for coding tasks.
  • Apple Silicon Users: Maximizing performance on macOS with MLX-accelerated inference.
  • Zed IDE Users: Direct integration with Zed's Agent Panel and Inline Assistant for tool-calling capabilities.
  • Privacy-Conscious Development: Keeping code generation and analysis entirely local without relying on external cloud services.