Jolly-Q/llma31_base_33_ST

TEXT GENERATIONConcurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Nov 6, 2025Architecture:Transformer Featherless Exclusive Cold

Jolly-Q/llma31_base_33_ST is a 70 billion parameter base model derived from Llama 3.1, featuring transplanted Llama 3.3 special token embeddings. This model is specifically designed as a foundation for instruction-tuned QLoRA fine-tuning, simplifying the process by eliminating the need to target embeddings or the language model head. It is not intended for direct chat inference but serves as an optimized base for developers to create specialized instruction-following models.

Loading preview...

Model Overview

Jolly-Q/llma31_base_33_ST is a 70 billion parameter base model built upon the Llama 3.1 architecture. Its key distinguishing feature is the integration of Llama 3.3 special token embeddings, which are transplanted into the Llama 3.1 base weights.

Key Capabilities

  • Optimized for Fine-tuning: This model is specifically engineered to facilitate instruction-tuned QLoRA fine-tuning. The transplanted embeddings streamline the process.
  • Simplified Tuning: Developers can perform QLoRA tuning without the complexity of targeting the embeddings or the lm_head directly, making the fine-tuning workflow more efficient.

Good For

  • Instruction-tuned QLoRA: Ideal for researchers and developers looking to create custom instruction-following models based on Llama 3.1 with enhanced tuning simplicity.
  • Base Model for Customization: Serves as a robust foundation for various downstream tasks requiring specialized instruction-following capabilities.

Important Note

This model is not suitable for chat inference in its base state. It is intended solely as a foundational model for further fine-tuning and development.