ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 20, 2026Architecture:Transformer Featherless Exclusive Cold

The ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth is a 4.0 billion parameter Qwen3ForCausalLM architecture, LoRA fine-tuned from Microsoft's FastContext-1.0-4B-RL. It was supervised fine-tuned on a private Claude reasoning distillation dataset, making it suitable for tasks requiring enhanced reasoning capabilities. This model leverages Unsloth and TRL for efficient training, focusing on instruction-following behavior.

Loading preview...

Model Overview

This model, ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth, is a 4.0 billion parameter language model built on the Qwen3ForCausalLM architecture. It is a LoRA fine-tune of Microsoft's FastContext-1.0-4B-RL base model.

Key Characteristics

  • Base Model: microsoft/FastContext-1.0-4B-RL
  • Architecture: Qwen3ForCausalLM with 4.0 billion parameters.
  • Fine-tuning: Supervised fine-tuned using LoRA (rank 16, alpha 16) via Unsloth and TRL.
  • Training Data: Fine-tuned on a private dataset, ermiaazarkhalili/claude-reasoning-distillation, specifically configured for supervised fine-tuning.
  • Precision: Trained with 4-bit QLoRA precision.

Limitations

  • No Benchmarks: No downstream benchmark evaluations have been conducted; only training loss observations are available.
  • Inherited Biases: The model inherits biases, knowledge cutoffs, and failure modes from its base model.
  • Untested Behavior: Fine-tuned on a single instruction-following dataset, its performance outside this distribution is untested.
  • Merged Adapters: LoRA adapters are merged into the base weights, preventing detachment from this specific fine-tune.