atenareply/gemma-4-12b-asterion

TEXT GENERATIONConcurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

atenareply/gemma-4-12b-asterion is a 12 billion parameter Gemma-4 model that has undergone full-parameter Continued Pre-Training (CPT) on the Asterion Space Operations corpus, a 1.887 billion token dataset. This model is specifically designed as a domain-knowledge backbone for fictional satellite operations, demonstrating significant perplexity reduction on both Asterion and Mars telemetry data. It is optimized for specialized domain understanding rather than general instruction following or chat capabilities.

Loading preview...

Overview of atenareply/gemma-4-12b-asterion

atenareply/gemma-4-12b-asterion is a 12 billion parameter Gemma-4 model that has undergone full-parameter Continued Pre-Training (CPT). This model was trained on a specialized 1.887 billion token corpus, primarily consisting of fictional Asterion Space Operations data (85%) and reused Mars Express telemetry (3%), with 12% FineWeb-Edu replay to prevent catastrophic forgetting. Training ceased at an early plateau after approximately 229 million tokens, indicating efficient domain knowledge acquisition.

Key Capabilities and Performance

  • Domain Specialization: Achieves a perplexity (PPL) of 1.83 on held-out Asterion data, a 62% improvement over the base Gemma-4-12B's 4.80 PPL. It also shows a PPL of 1.25 on Mars telemetry, down from 3.67.
  • Memory Efficiency: Demonstrates the ability to perform a 12B full-parameter fine-tune on a single H200 GPU, utilizing paged 8-bit AdamW and gradient checkpointing. This innovation addresses the memory constraints typically associated with large model training.
  • Anti-Forgetting: The inclusion of FineWeb-Edu replay during CPT successfully maintained general knowledge, with a PPL of 8.34 on general data, comparable to the base model's 8.55.

Intended Use and Limitations

This model serves as a domain-knowledge backbone for the fictional Asterion round, excelling in understanding and generating content related to invented satellite operations. It is not instruction-tuned and lacks inherent chat or tool-use capabilities, functioning as a base-style CPT checkpoint. Its knowledge is limited to the seen portion of the corpus, with specific per-document facts in the unseen 88% intended to be provided via downstream tasks or tool results in context.