nbeerbower/Hemlock-Jan-code-4B
nbeerbower/Hemlock-Jan-code-4B is a 4 billion parameter language model, fine-tuned from the janhq/Jan-code-4b base model using supervised fine-tuning (SFT). This model is optimized for code-related tasks, leveraging a 32768 token context length and 4-bit quantization for efficient deployment. It is designed for applications requiring robust code generation and understanding capabilities.
Loading preview...
Model Overview
nbeerbower/Hemlock-Jan-code-4B is a 4 billion parameter language model, developed by nbeerbower, that has been fine-tuned from the janhq/Jan-code-4b base model. This model utilizes a supervised fine-tuning (SFT) approach, making it suitable for specialized applications, particularly in code-related domains. It supports a substantial context length of 32768 tokens, allowing for processing longer code snippets and complex programming contexts.
Key Training Details
The fine-tuning process for Hemlock-Jan-code-4B involved specific configurations to enhance its performance:
- Base Model:
janhq/Jan-code-4b - Training Mode: Supervised Fine-Tuning (SFT)
- Quantization: 4-bit (NF4) for efficient memory usage and inference.
- Optimizer:
paged_adamw_8bitwith a cosine learning rate scheduler. - LoRA Configuration: Employed LoRA (Low-Rank Adaptation) with a rank of 128 and an alpha of 64, targeting key attention and feed-forward projection modules (
up_proj,down_proj,gate_proj,k_proj,q_proj,v_proj,o_proj).
Use Cases
Given its origin from a code-focused base model and SFT training, nbeerbower/Hemlock-Jan-code-4B is well-suited for:
- Code Generation: Assisting developers in writing new code or completing existing functions.
- Code Understanding: Analyzing and explaining code logic.
- Code Refactoring: Suggesting improvements or alternative implementations for code segments.
- Educational Tools: Providing programming assistance in learning environments.