yingfanbot/gsm-cot-llama1b
The yingfanbot/gsm-cot-llama1b is a 1 billion parameter Llama-3.2-1B-Instruct model, developed by Ying Fan, Anej Svete, and Kangwook Lee, that has been supervised fine-tuned for Chain-of-Thought (CoT) reasoning on the GSM8K dataset. This model serves as a baseline for the LOTUS (Looped Transformers with parallel supervision on latents) framework. It is specifically optimized for mathematical reasoning tasks, demonstrating enhanced performance in generating step-by-step solutions.
Loading preview...
Model Overview
The yingfanbot/gsm-cot-llama1b is a specialized language model built upon the meta-llama/Llama-3.2-1B-Instruct architecture. Developed by Ying Fan, Anej Svete, and Kangwook Lee, this model has undergone supervised fine-tuning (SFT) specifically for Chain-of-Thought (CoT) reasoning on the GSM8K dataset. This fine-tuning process enhances its ability to generate detailed, step-by-step solutions for mathematical word problems.
Key Capabilities
- Chain-of-Thought Reasoning: Excels at breaking down complex problems into logical, sequential steps.
- Mathematical Problem Solving: Optimized for arithmetic and reasoning tasks found in datasets like GSM8K.
- LOTUS Baseline: Serves as the Stage 1 SFT baseline for the LOTUS (Looped Transformers with parallel supervision on latents) framework.
Good For
- Research in Reasoning: Ideal for researchers exploring CoT mechanisms and latent reasoning in LLMs.
- Educational Tools: Can be integrated into applications requiring step-by-step explanations for math problems.
- Benchmarking: Useful as a baseline for evaluating new methods in mathematical reasoning and CoT generation.