XinnanZhang/Alfworld-qwen2.5-3b-SFT

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 29, 2025Architecture:Transformer Featherless Exclusive Cold

XinnanZhang/Alfworld-qwen2.5-3b-SFT is a 3.1 billion parameter language model based on the Qwen2.5 architecture. This model is a fine-tuned version, indicated by 'SFT', suggesting it has undergone Supervised Fine-Tuning for specific tasks. With a context length of 32768 tokens, it is designed for applications requiring processing of longer sequences. Its specific differentiators and primary use cases are not detailed in the provided model card, which indicates 'More Information Needed' for most sections.

Loading preview...

Model Overview

This model, XinnanZhang/Alfworld-qwen2.5-3b-SFT, is a 3.1 billion parameter language model built upon the Qwen2.5 architecture. The 'SFT' in its name indicates that it has undergone Supervised Fine-Tuning, typically to adapt a base model for specific downstream tasks or improved instruction following. It supports a substantial context length of 32768 tokens, allowing it to process and generate longer text sequences.

Key Characteristics

  • Architecture: Based on the Qwen2.5 model family.
  • Parameter Count: 3.1 billion parameters, making it a relatively compact yet capable model.
  • Context Window: Features a large context window of 32768 tokens, beneficial for tasks requiring extensive contextual understanding or generation.
  • Fine-Tuning: The 'SFT' designation implies it has been fine-tuned, likely for enhanced performance on particular tasks, though the specific training data or objectives are not detailed in the current model card.

Current Limitations

The provided model card indicates that significant details regarding its development, specific use cases, training data, evaluation metrics, and potential biases are currently marked as "More Information Needed." Users should be aware that without further documentation, the precise capabilities, intended applications, and limitations of this specific fine-tuned model are not fully disclosed.