Neelectric/Llama-3.1-8B-Instruct_SFT_mathsp_ewc_v00.15_s44

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026Architecture:Transformer Featherless Exclusive Cold

Neelectric/Llama-3.1-8B-Instruct_SFT_mathsp_ewc_v00.15_s44 is an 8 billion parameter instruction-tuned language model, fine-tuned from Meta's Llama-3.1-8B-Instruct. It was specifically trained on the Neelectric/OpenR1-Math-220k_all_Llama3_4096toks dataset, making it optimized for mathematical reasoning and problem-solving tasks. With a 32768 token context length, this model is well-suited for applications requiring robust mathematical capabilities.

Loading preview...

Model Overview

This model, Neelectric/Llama-3.1-8B-Instruct_SFT_mathsp_ewc_v00.15_s44, is an 8 billion parameter instruction-tuned large language model. It is a specialized version of the meta-llama/Llama-3.1-8B-Instruct base model, fine-tuned by Neelectric.

Key Capabilities

  • Mathematical Reasoning: The model has undergone supervised fine-tuning (SFT) on the Neelectric/OpenR1-Math-220k_all_Llama3_4096toks dataset, indicating a strong focus on mathematical problem-solving and reasoning tasks.
  • Instruction Following: As an instruction-tuned model, it is designed to follow user prompts and generate relevant responses effectively.
  • Extended Context: It supports a substantial context length of 32768 tokens, allowing for processing and understanding longer inputs and complex problem descriptions.

Training Details

The model was trained using the TRL library for supervised fine-tuning. The training process utilized specific versions of popular ML frameworks, including TRL 1.1.0.dev0, Transformers 4.57.6, Pytorch 2.9.0, Datasets 5.0.1, and Tokenizers 0.22.2.

Good For

  • Applications requiring strong mathematical understanding and problem-solving.
  • Tasks involving complex instructions where a robust instruction-following model is beneficial.
  • Scenarios needing a model capable of handling long input contexts, particularly in technical or analytical domains.