trashpanda-org/gemma-4-31b-larkspur-v1
VISIONPricing:Input $0.48 / Cached $0.1 / Output $1.44Concurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 26, 2026Architecture:Transformer0.0K Featherless Exclusive Cold
The trashpanda-org/gemma-4-31b-larkspur-v1 is a 31 billion parameter language model based on the Gemma-4 architecture, developed by trashpanda-org. This model utilizes r64a32 rslora and has liger_use_token_scaling enabled, distinguishing it from its predecessor. It is designed for general language understanding and generation tasks, offering a substantial parameter count for complex applications.
Loading preview...
Overview
The trashpanda-org/gemma-4-31b-larkspur-v1 is a 31 billion parameter language model built upon the Gemma-4 architecture. Developed by trashpanda-org, this version represents an iteration over its predecessor, v0, by incorporating specific architectural and training modifications.
Key Characteristics
- Parameter Count: Features 31 billion parameters, providing a robust foundation for various natural language processing tasks.
- Architecture: Based on the Gemma-4 model family, known for its efficiency and performance.
- LoRA Configuration: Employs
r64a32 rslorainstead of ther32a16configuration found inv0, suggesting a different approach to low-rank adaptation during fine-tuning. - Token Scaling: Integrates
liger_use_token_scaling, which is enabled in this version. This feature likely influences how token representations are handled and scaled during the model's operation, potentially improving performance or efficiency. - Dataset Consistency: Utilizes the same dataset as its
v0counterpart, indicating that the improvements inlarkspur-v1stem primarily from architectural and training methodology changes rather than new data.
Good For
- Applications requiring a large language model with 31 billion parameters.
- Exploring the impact of
r64a32 rsloraandliger_use_token_scalingon Gemma-4 based models. - General-purpose language generation and understanding tasks where the Gemma-4 architecture is preferred.