sagnikM/qwen_qwen_step75_posterior

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026Architecture:Transformer Featherless Exclusive Cold

The sagnikM/qwen_qwen_step75_posterior is a 7.6 billion parameter Qwen2.5-7B causal language model, developed by sagnikM, functioning as a privileged hint generator. This model is derived from the step-75 FSDP checkpoint of the HiLL Qwen2.5-7B/OpenThoughts run. It is specifically designed for tasks requiring hint generation, leveraging its 32768 token context length.

Loading preview...

Model Overview

The sagnikM/qwen_qwen_step75_posterior is a 7.6 billion parameter causal language model based on the Qwen2.5-7B architecture. It serves as a privileged hint generator, specifically converted from the step-75 FSDP checkpoint of the HiLL Qwen2.5-7B/OpenThoughts training run. This model is designed to provide hints, likely in a specialized context related to its training origin.

Key Characteristics

  • Architecture: Qwen2.5-7B, a causal language model.
  • Parameter Count: 7.6 billion parameters.
  • Context Length: Supports a context window of 32768 tokens.
  • Origin: Derived from a specific training step (step 75) of the HiLL Qwen2.5-7B/OpenThoughts project.
  • Functionality: Primarily acts as a privileged hint generator.

Usage Notes

This repository contains only the model, configuration, and tokenizer artifacts. The optimizer and trainer states from the original VERL checkpoint are not included. Developers can load the model and tokenizer using the transformers library for inference tasks related to hint generation.