ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-es

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-es model is a fine-tuned version of the Qwen3-1.7B architecture, featuring 2 billion parameters and a 32768-token context length. This model has been specifically fine-tuned over 3 epochs with a focus on reasoning capabilities, as indicated by 'reas-int-065'. It is designed for tasks requiring robust inference and understanding, making it suitable for applications where logical processing is crucial.

Loading preview...

Model Overview

This model, ferrazzipietro/Qwen3-1.7B-reas-int-065-3-epochs-es, is a fine-tuned variant of the Qwen3-1.7B base model, developed by Qwen. It incorporates 2 billion parameters and supports a substantial context length of 32768 tokens, making it capable of processing extensive inputs.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen3-1.7B model.
  • Fine-tuning Focus: The model name suggests a specialization in 'reasoning' (reas-int-065), indicating an optimization for tasks that require logical inference and problem-solving.
  • Training Details: Fine-tuned over 3 epochs using specific hyperparameters including a learning rate of 5e-06, a total batch size of 32, and an AdamW optimizer. The training utilized a cosine learning rate scheduler with a 0.1 warmup ratio.

Potential Use Cases

Given its fine-tuning for reasoning, this model is likely well-suited for:

  • Complex Question Answering: Handling queries that require multi-step logic or inference.
  • Textual Analysis: Tasks involving understanding relationships, implications, or causality within text.
  • Problem Solving: Applications where the model needs to deduce solutions from given information.

Further details on the specific dataset used for fine-tuning and comprehensive evaluation results are not provided in the current documentation.