ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 20, 2026Architecture:Transformer Featherless Exclusive Cold
The ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth is a 4.0 billion parameter Qwen3ForCausalLM architecture, LoRA fine-tuned from Microsoft's FastContext-1.0-4B-RL. It was supervised fine-tuned on a private Claude reasoning distillation dataset, making it suitable for tasks requiring enhanced reasoning capabilities. This model leverages Unsloth and TRL for efficient training, focusing on instruction-following behavior.
Loading preview...
Model Overview
This model, ermiaazarkhalili/FastContext-4B-RL_base-SFT-Claude-Opus-Reasoning-Unsloth, is a 4.0 billion parameter language model built on the Qwen3ForCausalLM architecture. It is a LoRA fine-tune of Microsoft's FastContext-1.0-4B-RL base model.
Key Characteristics
- Base Model:
microsoft/FastContext-1.0-4B-RL - Architecture:
Qwen3ForCausalLMwith 4.0 billion parameters. - Fine-tuning: Supervised fine-tuned using LoRA (rank 16, alpha 16) via Unsloth and TRL.
- Training Data: Fine-tuned on a private dataset,
ermiaazarkhalili/claude-reasoning-distillation, specifically configured for supervised fine-tuning. - Precision: Trained with 4-bit QLoRA precision.
Limitations
- No Benchmarks: No downstream benchmark evaluations have been conducted; only training loss observations are available.
- Inherited Biases: The model inherits biases, knowledge cutoffs, and failure modes from its base model.
- Untested Behavior: Fine-tuned on a single instruction-following dataset, its performance outside this distribution is untested.
- Merged Adapters: LoRA adapters are merged into the base weights, preventing detachment from this specific fine-tune.