g-assismoraes/Qwen3-4B-Instruct-2507-imdb

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 6, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

g-assismoraes/Qwen3-4B-Instruct-2507-imdb is a 4 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen3-4B-Instruct-2507. This model is optimized for specific tasks, demonstrating a validation loss of 1.6597. It is suitable for applications requiring a compact yet capable language model for fine-tuned instruction following.

Loading preview...

Model Overview

This model, g-assismoraes/Qwen3-4B-Instruct-2507-imdb, is a fine-tuned variant of the Qwen3-4B-Instruct-2507 base model. It features 4 billion parameters and is designed for instruction-following tasks. The fine-tuning process resulted in a validation loss of 1.6597, indicating its performance on the specific dataset it was trained on.

Key Characteristics

  • Base Model: Qwen/Qwen3-4B-Instruct-2507
  • Parameter Count: 4 billion
  • Context Length: 32768 tokens
  • Training Objective: Instruction-tuning, with a focus on achieving a low validation loss.

Training Details

The model was trained with a learning rate of 2e-05 over 2 epochs, using a batch size of 4 for both training and evaluation. The optimizer used was ADAMW_TORCH. The training process showed a decrease in loss from 1.6104 in the first epoch to 1.5466 in the second, with the final validation loss recorded at 1.6597.

Potential Use Cases

Given its instruction-tuned nature and specific fine-tuning, this model is likely suitable for:

  • Specific domain applications where the fine-tuning dataset aligns with the task.
  • Instruction-following tasks requiring a compact 4B parameter model.
  • Research and experimentation with fine-tuned Qwen3 models.