SirSahOl/gemma-3-1b-it-chat-mlx-16bit
SirSahOl/gemma-3-1b-it-chat-mlx-16bit is a 1 billion parameter Gemma3ForCausalLM model, a 16-bit MLX conversion of Google's gemma-3-1b-it. Optimized for Apple Silicon native GPU inference, it offers full unquantized precision with a 32,768 token context length. This model is designed for high-end workstation deployments, research, and benchmarking on M2/M3/M4 Max/Ultra chips with 32GB+ unified memory.
Loading preview...
Overview
SirSahOl/gemma-3-1b-it-chat-mlx-16bit is a 1 billion parameter Gemma3ForCausalLM model, specifically a 16-bit MLX conversion of the original google/gemma-3-1b-it. This conversion is optimized for native GPU inference on Apple Silicon (M1/M2/M3/M4 series) using Apple's MLX framework. It provides full unquantized precision, making it suitable for tasks requiring the highest possible accuracy.
Key Capabilities & Features
- Architecture: Gemma3ForCausalLM with 1 billion parameters.
- Context Length: Supports a substantial context window of 32,768 tokens.
- Precision: 16-bit (bfloat16) quantization, offering reference evaluation quality with no quality degradation.
- Hardware Optimization: Tailored for Apple Silicon, leveraging MLX for efficient on-device inference.
- VRAM Footprint: Requires approximately 2.8 GB of VRAM, with a minimum recommended 8 GB Unified Memory.
Performance & Hardware Considerations
On an Apple M1, this 16-bit variant achieves 20.92 tokens/sec, with a Time To First Token (TTFT) of 48.02 ms. While offering the highest precision, it has a lower token generation speed compared to its 4-bit and 8-bit counterparts. It is best suited for high-end Apple Silicon hardware like M2/M3/M4 Max/Ultra with 32GB+ unified memory.
Ideal Use Cases
- Research and Benchmarking: Provides unquantized full precision for accurate evaluation and research.
- High-End Workstation Deployments: Suitable for environments where maximum quality is prioritized over raw speed.
- Precision-Critical Reasoning: Recommended for tasks where minor precision loss from lower quantizations is unacceptable.