upstage/SOLAR-10.7B-v1.0

Hugging Face
TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:10.7BQuant:FP8Context Size:4kPublished:Dec 12, 2023License:apache-2.0Architecture:Transformer0.3K Open Weights Featherless Exclusive Warm

SOLAR-10.7B-v1.0 by Upstage is a 10.7 billion parameter large language model, developed using a Depth Up-Scaling (DUS) methodology that integrates Mistral 7B weights. This compact yet powerful model demonstrates superior performance, outperforming many models up to 30B parameters, including Mixtral 8x7B, on the H6 benchmark. It is primarily designed as a robust and adaptable base model for fine-tuning, offering significant performance improvements for various natural language processing tasks.

Loading preview...

SOLAR-10.7B-v1.0: Compact Powerhouse by Upstage

SOLAR-10.7B-v1.0 is a 10.7 billion parameter large language model developed by Upstage. It introduces a novel Depth Up-Scaling (DUS) methodology, which involves integrating Mistral 7B weights into upscaled layers followed by continued pre-training. This approach results in a remarkably powerful model despite its compact size.

Key Capabilities & Performance

  • Superior Performance: SOLAR-10.7B-v1.0 demonstrates exceptional performance, surpassing many larger models, including the Mixtral 8x7B, on the H6 evaluation benchmark. The instruction-tuned variant, SOLAR-10.7B-Instruct-v1.0, achieves a score of 74.20 on H6, outperforming Mixtral-8x7B-Instruct-v0.1 (72.62).
  • Robust Base Model: It is specifically highlighted as an ideal choice for fine-tuning, offering high robustness and adaptability for various downstream NLP tasks.
  • Efficient Scaling: The DUS methodology allows for significant performance gains without a proportional increase in model size, making it a highly efficient LLM.

Ideal Use Cases

  • Fine-tuning: This model is primarily intended as a strong pre-trained base for further instruction fine-tuning to adapt it to specific applications and conversational agents.
  • Research & Development: Developers and researchers looking for a powerful yet compact LLM for experimentation and building custom solutions will find its performance and scaling methodology valuable.
  • Resource-Constrained Environments: Its relatively smaller size compared to models it outperforms makes it suitable for scenarios where computational resources are a consideration.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p