AtlaAI/Selene-1-Mini-Llama-3.1-8B

Hugging Face
TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 22, 2025License:apache-2.0Architecture:Transformer0.1K Open Weights Featherless Exclusive Warm

AtlaAI/Selene-1-Mini-Llama-3.1-8B is an 8 billion parameter small language model-as-a-judge (SLMJ) developed by Atla. Post-trained from Llama-3.1-8B, it excels in evaluation tasks, outperforming larger models like GPT-4o on RewardBench, EvalBiasBench, and AutoJ. This model is optimized for general-purpose evaluation, supporting absolute scoring, classification, and pairwise preference tasks across a 128K context length.

Loading preview...

Model Overview

Atla Selene Mini is an 8 billion parameter small language model-as-a-judge (SLMJ) developed by Atla. It is post-trained from Llama-3.1-8B and designed for robust evaluation tasks. The model demonstrates performance comparable to models 10x its size, notably outperforming GPT-4o on RewardBench, EvalBiasBench, and AutoJ.

Key Capabilities

  • Advanced Evaluation: Excels across 11 benchmarks covering absolute scoring (e.g., harmlessness on a 1-5 scale), classification (e.g., addressing user queries with Yes/No), and pairwise preference (e.g., logical consistency comparison).
  • Top Performance: Ranks as the #1 8B generative model on RewardBench.
  • Multilingual Support: Primarily English, but also supports German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
  • Structured Outputs: Generates structured evaluation outputs and provides qualitative critiques with reasoning.
  • Extended Context: Features a 128K context length, enabling comprehensive analysis of longer inputs.

Good For

  • General-purpose evaluation: Ideal for assessing model responses, agent performance, and content quality.
  • Absolute scoring: Evaluating responses based on specific criteria and scales.
  • Classification tasks: Determining if responses meet predefined conditions.
  • Pairwise preference: Comparing and ranking responses based on desired attributes.
  • RAG hallucination detection: Cookbooks are provided for specific use cases like RAG hallucination evaluation.

Users should apply the Llama 3 conversation template for optimal performance, with prompt templates used during training available in the cookbooks.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p