talzoomanzoo/qwen2_5_3b_uid_reference_lr3e6_step2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The talzoomanzoo/qwen2_5_3b_uid_reference_lr3e6_step2 is a 3.1 billion parameter causal language model, derived from Qwen/Qwen2.5-3B. This model incorporates a merged LoRA adapter from a specific training run, making it a fine-tuned variant. It is designed for direct loading and use without additional adapters, offering a specialized iteration of the Qwen2.5-3B base model.

Loading preview...

Model Overview

The talzoomanzoo/qwen2_5_3b_uid_reference_lr3e6_step2 is a specialized variant of the Qwen/Qwen2.5-3B causal language model. This model integrates the full weights from a specific training run, where a LoRA (Low-Rank Adaptation) actor from step 2 was merged into the base Qwen2.5-3B model.

Key Characteristics

  • Base Model: Qwen/Qwen2.5-3B
  • Parameter Count: 3.1 billion parameters
  • Context Length: 32768 tokens
  • Integration: The model includes a merged LoRA adapter, meaning no separate adapter is required for use.
  • Training Details: The merged LoRA was part of a training job (ID: 3977891) at global step 2, utilizing a learning rate of 3e-6, with a LoRA rank of 32 and alpha of 16.
  • Export Dtype: The model weights are exported in bfloat16 format.

Usage

This model can be loaded directly using AutoModelForCausalLM.from_pretrained, simplifying its integration into existing workflows. It represents a specific fine-tuned state of the Qwen2.5-3B architecture, optimized through the described LoRA training process.