trashpanda-org/gemma-4-31b-larkspur-v1

VISIONPricing:Input $0.48 / Cached $0.1 / Output $1.44Concurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 26, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

The trashpanda-org/gemma-4-31b-larkspur-v1 is a 31 billion parameter language model based on the Gemma-4 architecture, developed by trashpanda-org. This model utilizes r64a32 rslora and has liger_use_token_scaling enabled, distinguishing it from its predecessor. It is designed for general language understanding and generation tasks, offering a substantial parameter count for complex applications.

Loading preview...

Overview

The trashpanda-org/gemma-4-31b-larkspur-v1 is a 31 billion parameter language model built upon the Gemma-4 architecture. Developed by trashpanda-org, this version represents an iteration over its predecessor, v0, by incorporating specific architectural and training modifications.

Key Characteristics

  • Parameter Count: Features 31 billion parameters, providing a robust foundation for various natural language processing tasks.
  • Architecture: Based on the Gemma-4 model family, known for its efficiency and performance.
  • LoRA Configuration: Employs r64a32 rslora instead of the r32a16 configuration found in v0, suggesting a different approach to low-rank adaptation during fine-tuning.
  • Token Scaling: Integrates liger_use_token_scaling, which is enabled in this version. This feature likely influences how token representations are handled and scaled during the model's operation, potentially improving performance or efficiency.
  • Dataset Consistency: Utilizes the same dataset as its v0 counterpart, indicating that the improvements in larkspur-v1 stem primarily from architectural and training methodology changes rather than new data.

Good For

  • Applications requiring a large language model with 31 billion parameters.
  • Exploring the impact of r64a32 rslora and liger_use_token_scaling on Gemma-4 based models.
  • General-purpose language generation and understanding tasks where the Gemma-4 architecture is preferred.