yuq-zhou/2026-05-o-b0p3-a1p0-gc0p75-exp-td4p0-tw10p0-mbz-r1-7-last

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026Architecture:Transformer Featherless Exclusive Cold

The yuq-zhou/2026-05-o-b0p3-a1p0-gc0p75-exp-td4p0-tw10p0-mbz-r1-7-last model is a 7.6 billion parameter language model with a 32K context length, provided in standard HuggingFace format. This model is a research artifact backup, primarily intended for internal or specific research purposes. Its main utility lies in serving as a checkpoint for further experimentation or analysis within a research context.

Loading preview...

Overview

The yuq-zhou/2026-05-o-b0p3-a1p0-gc0p75-exp-td4p0-tw10p0-mbz-r1-7-last is a 7.6 billion parameter language model with a 32,768 token context length. It is provided as a model checkpoint in the standard HuggingFace format, making it compatible with AutoModelForCausalLM.from_pretrained for easy loading and use within the HuggingFace ecosystem.

Key Characteristics

  • Parameter Count: 7.6 billion parameters, indicating a moderately large model capable of complex language understanding and generation tasks.
  • Context Length: A substantial context window of 32,768 tokens, allowing it to process and generate long sequences of text while maintaining coherence and understanding.
  • Format: Distributed in the widely adopted HuggingFace format, ensuring broad compatibility and ease of integration into existing ML workflows.

Intended Use

This model is primarily designated as a research artifact backup. Its main purpose is to serve as a stable checkpoint for research and development, rather than a general-purpose, instruction-tuned model for broad applications. It is suitable for:

  • Research and Experimentation: Researchers can use this checkpoint for further fine-tuning, architectural analysis, or as a baseline for new experiments.
  • Internal Development: Teams working on specific projects might use this as a foundational model for specialized tasks.

Given its nature as a research artifact, users should be aware that it may not be optimized for direct end-user applications without further fine-tuning or instruction-tuning.