xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL
The xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL model is a 3.1 billion parameter instruction-tuned language model based on the Qwen2.5 architecture. This model is specifically fine-tuned with a focus on MMLU performance, utilizing a GRPO_KL training regime. It is designed for general instruction-following tasks, leveraging a 32768 token context length. The model aims to provide robust performance in academic and reasoning benchmarks within its parameter class.
Loading preview...
Model Overview
This model, xw1234gan/GRPO_KL_Qwen2.5-3B-Instruct_MMLU_beta0_lr1e-05_mb2_ga128_n2048_seed42_NoKL, is an instruction-tuned variant of the Qwen2.5-3B architecture, featuring 3.1 billion parameters and a substantial 32768 token context window. It has been developed by xw1234gan with a specific emphasis on optimizing performance on the MMLU (Massive Multitask Language Understanding) benchmark, indicated by its GRPO_KL training methodology and MMLU-focused naming convention.
Key Characteristics
- Architecture: Based on the Qwen2.5-3B model.
- Parameter Count: 3.1 billion parameters, offering a balance between performance and computational efficiency.
- Context Length: Supports a long context window of 32768 tokens, enabling processing of extensive inputs.
- Training Focus: Fine-tuned using a GRPO_KL regime with specific parameters (beta0, learning rate, batch size, gradient accumulation, and seed) to enhance MMLU performance.
Potential Use Cases
Given its instruction-tuned nature and MMLU optimization, this model is suitable for:
- General Instruction Following: Responding to a wide array of user prompts and instructions.
- Academic and Reasoning Tasks: Performing well on tasks requiring understanding and reasoning, as suggested by its MMLU focus.
- Applications requiring long context: Handling detailed documents, conversations, or code with its extended context window.