Neelectric/Llama-3.1-8B-Instruct_SFT_mathsp_ewc_v00.15_s43
Neelectric/Llama-3.1-8B-Instruct_SFT_mathsp_ewc_v00.15_s43 is an 8 billion parameter instruction-tuned language model developed by Neelectric, fine-tuned from Meta's Llama-3.1-8B-Instruct. This model specializes in mathematical reasoning and problem-solving, having been trained on a dedicated mathematical dataset. With a 32,768 token context length, it is optimized for complex mathematical tasks and detailed analytical responses.
Loading preview...
Overview
This model, developed by Neelectric, is a specialized instruction-tuned variant of Meta's Llama-3.1-8B-Instruct, featuring 8 billion parameters and a 32,768 token context window. It has been specifically fine-tuned using Supervised Fine-Tuning (SFT) on the Neelectric/OpenR1-Math-220k_all_Llama3_4096toks dataset. This targeted training aims to enhance its capabilities in mathematical reasoning and problem-solving.
Key Capabilities
- Enhanced Mathematical Reasoning: Optimized for understanding and generating responses to mathematical queries.
- Instruction Following: Retains strong instruction-following abilities from its base Llama-3.1-8B-Instruct model.
- Extended Context: Benefits from a 32,768 token context length, allowing for processing longer and more complex mathematical problems or discussions.
Good For
- Applications requiring robust mathematical problem-solving.
- Educational tools focused on math assistance.
- Research in mathematical language understanding and generation.
Training Details
The model was trained using the TRL library (version 1.1.0.dev0) with PyTorch 2.9.0, Transformers 4.57.6, and Datasets 5.0.1. The training process involved SFT on a comprehensive mathematical dataset, indicating a strong focus on numerical and logical tasks.