aqbond/deepswe-k8s-sync-S0b16-c256-step140to180-step175

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 31, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The aqbond/deepswe-k8s-sync-S0b16-c256-step140to180-step175 model is a 32 billion parameter language model, specifically a FSDP-merged Hugging Face weight checkpoint from training step 175. This model is derived from a deepswe-k8s-sync PPO recipe without trajectory repair. It represents a specific snapshot of a training experiment, making it suitable for research and evaluation of the deepswe-k8s-sync training methodology.

Loading preview...

Model Overview

This model, aqbond/deepswe-k8s-sync-S0b16-c256-step140to180-step175, is a 32 billion parameter checkpoint representing the state of a training run at step 175. It consists of FSDP-merged Hugging Face weights, indicating its origin from a distributed training setup.

Key Characteristics

  • Training Snapshot: This model is a specific checkpoint from training step 175 of the deepswe-k8s-sync-S0b16-c256-step140to180-20260728-231847 experiment.
  • Training Recipe: It was trained using a S0b16 sync PPO (Proximal Policy Optimization) recipe, notably without trajectory repair.
  • Checkpoint Source: The weights were sourced from a specific actor checkpoint within the training outputs.
  • Paired Repository: It has a corresponding replacement repository, aqbond/deepswe-k8s-sync-S0b16-c256-replace-tokenexact-step175, which might be relevant for specific evaluation or fine-tuning tasks.

Good For

  • Research and Development: Ideal for researchers and developers interested in evaluating the performance and characteristics of models trained with the deepswe-k8s-sync PPO recipe at a specific training stage.
  • Comparative Analysis: Useful for comparing the evolution of model capabilities across different training steps or against models trained with alternative PPO configurations.
  • Further Fine-tuning: Can serve as a base model for continued fine-tuning or experimentation, particularly for tasks related to the original training objective of the deepswe-k8s-sync experiment.