youngseok12/AX-4.0-Light-sft_71949_text_causal

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The youngseok12/AX-4.0-Light-sft_71949_text_causal is a 7.6 billion parameter Qwen2-family causal language model, derived from skt/A.X-4.0-Light, with a 32768 token context length. It was fine-tuned using a LoRA adapter on a filtered, text-only subset of the Korean AI Hub Dataset 71949, specifically for causal reasoning tasks in Korean. This model is designed for Korean-language research and controlled benchmark experiments, focusing on improving reasoning capabilities.

Loading preview...

Model Overview

This model, youngseok12/AX-4.0-Light-sft_71949_text_causal, is a 7.6 billion parameter Qwen2-family causal language model based on skt/A.X-4.0-Light. It has been fine-tuned using a LoRA adapter (rank 16, alpha 32) on a specific subset of the Korean AI Hub Dataset 71949, focusing exclusively on causal reasoning text data. The LoRA adapter was merged into the base weights, resulting in a standalone BF16 full-weight model.

Key Capabilities & Training

  • Specialized Fine-tuning: Trained on a carefully filtered, text-only subset of AI Hub Dataset 71949, which contains causal reasoning examples in Korean. Image-grounded labels were excluded, and visual wording was normalized to text.
  • Architecture: Utilizes the Qwen2-family causal language model architecture, with the base architecture remaining unchanged.
  • Training Details: Trained with 478 examples over 2 epochs, using a sequence length of 2048 and BF16 precision.
  • Korean Language Focus: Primarily intended for Korean-language research and controlled benchmark experiments.

Performance Highlights

Local evaluation using deterministic probes yielded the following parsed accuracies:

  • KMMLU-Pro: 47.27%
  • CLIcK: 67.37%
  • SNU Ko-MuSR: 48.93%
  • Com2-main: 50.60%
    The five-axis local mean accuracy is 43.59%.

Usage & Limitations

This merged model can be loaded directly with standard Transformers or vLLM without requiring trust_remote_code or a separate adapter. It is an experimental model for research and evaluation, and may produce factual or reasoning errors. It is not suitable for professional advice systems.