ynklab/Qwen2.5-7B-Stair_2c1t

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ynklab/Qwen2.5-7B-Stair_2c1t is a 7.6 billion parameter language model developed by ynklab, fine-tuned from Qwen/Qwen2.5-7B-Instruct. It specializes in multilingual chunk-level machine translation, supporting English to and from German, Spanish, French, Italian, Korean, Dutch, Portuguese, Russian, and Chinese. This model utilizes a unique stair-step context scheme, processing translation chunks with zero, one, or two preceding source-language context chunks to maintain document-level consistency. It is optimized for document-level machine translation tasks, particularly for fixed-range chunking scenarios.

Loading preview...

Overview

ynklab/Qwen2.5-7B-Stair_2c1t is a 7.6 billion parameter model, fine-tuned from Qwen/Qwen2.5-7B-Instruct, specifically designed for multilingual chunk-level machine translation. It was developed as part of the "Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking" research, with code and inference scripts available on the Doc2FRC GitHub repository.

Key Capabilities

  • Multilingual Translation: Supports translation between English and German, Spanish, French, Italian, Korean, Dutch, Portuguese, Russian, and Chinese in both directions.
  • Context-Aware Chunking: Implements a unique Stair_2c1t variant, using a stair-step context scheme where translation of a current chunk can leverage zero, one, or two preceding source-language context chunks.
  • Document-Level Consistency: Fine-tuned on fixed-range chunks (256–512 tokens) derived from the sardinelab/DocBlocks dataset, aiming for improved consistency in document translation.
  • High Context Length: Supports a maximum sequence length of 32,768 tokens, enabling processing of longer document segments.

Training Details

The model underwent full-parameter supervised fine-tuning for 2 epochs, utilizing a learning rate of 7e-6 with a cosine scheduler and AdamW optimizer. It was trained with bfloat16 precision.

Recommended Use Cases

This model is ideal for developers and researchers focused on:

  • Document-level machine translation: Especially for scenarios requiring context preservation across document chunks.
  • Translating long texts: By breaking them into fixed-range chunks and leveraging the model's context handling.
  • Research in machine translation: Particularly for exploring chunk-level and document-level translation strategies.