positron-ai/Llama-3.1-8B-Instruct

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

Llama-3.1-8B-Instruct is an 8 billion parameter instruction-tuned causal language model developed by Meta. This model is an unmodified redistribution of the original Llama 3.1 series, featuring a 32,768 token context length. It is designed for general-purpose conversational AI and instruction following, serving as a foundational model for various NLP applications. This specific version is provided by Positron AI as a pinned source for their GPTQ quantizations.

Loading preview...

Overview

This model, Llama-3.1-8B-Instruct, is an unmodified redistribution of Meta's meta-llama/Llama-3.1-8B-Instruct. It is an 8 billion parameter instruction-tuned causal language model, part of the Llama 3.1 series, and features a substantial 32,768 token context length. Positron AI provides this snapshot as a stable, byte-identical source for their published GPTQ quantizations.

Key Characteristics

  • Architecture: Llama 3.1 series, a powerful causal language model.
  • Parameters: 8 billion, offering a balance of performance and efficiency.
  • Context Length: 32,768 tokens, enabling processing of extensive inputs and generating longer, more coherent responses.
  • Instruction-Tuned: Optimized for following instructions and engaging in conversational AI.
  • Provenance: Byte-identical to the original Meta release, ensuring consistency and reliability.

Usage and Licensing

Use of this model is governed by the Llama 3.1 Community License Agreement and Meta's Acceptable Use Policy. For comprehensive documentation, evaluations, and responsible-use guidance, users should refer to Meta's original model card.