espressovi/SUTRA-qwen3-8b-rlvr-close

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The espressovi/SUTRA-qwen3-8b-rlvr-close model is an 8 billion parameter language model based on the Qwen3-8B-Base architecture, fine-tuned using Reinforcement Learning (RL). With a context length of 32768 tokens, this model is specifically developed as an artifact for the SUTRA project. Its RL training distinguishes it from base models, suggesting optimizations for specific task performance or alignment.

Loading preview...

Model Overview

The espressovi/SUTRA-qwen3-8b-rlvr-close is an 8 billion parameter language model built upon the robust Qwen3-8B-Base architecture. This model is a key artifact of the SUTRA project, distinguished by its application of Reinforcement Learning (RL) during its training process.

Key Characteristics

  • Base Architecture: Leverages the foundational capabilities of the Qwen3-8B-Base model.
  • Parameter Count: Features 8 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended outputs.
  • Reinforcement Learning (RL) Fine-tuning: The "rlvr" in its name signifies that it has undergone Reinforcement Learning, which typically aims to align the model's outputs more closely with human preferences or specific task objectives, potentially improving aspects like instruction following, safety, or factual accuracy compared to its base counterpart.

Potential Use Cases

Given its RL-trained nature, this model is likely optimized for:

  • Specific Task Performance: Ideal for applications where the base Qwen3-8B model's outputs needed further refinement or alignment.
  • Research and Development: Suitable for researchers exploring the impact of RL fine-tuning on large language models, particularly within the context of the SUTRA project's goals.
  • Instruction Following: RL training often enhances a model's ability to follow complex instructions and generate desired response formats.