JackHsieh/qwen3-4B-instruct-luna-distill-step2362

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

JackHsieh/qwen3-4B-instruct-luna-distill-step2362 is a 4 billion parameter instruction-tuned Qwen3 model, distilled by JackHsieh. This model is specifically fine-tuned on 'gpt-5.6-luna' thoughts regarding the next 8 Qwen3 tokens of stat.ML arXiv LaTeX. It is optimized for generating highly specific, technical text related to machine learning research papers, achieving a final validation log probability of -0.9238 nats/token on held-out luna thoughts.

Loading preview...

Model Overview

JackHsieh/qwen3-4B-instruct-luna-distill-step2362 is a specialized 4 billion parameter instruction-tuned model based on the Qwen3 architecture. Developed by JackHsieh, this model has undergone a unique distillation process, leveraging 'gpt-5.6-luna' generated thoughts.

Key Capabilities

  • Specialized Text Generation: Fine-tuned to predict the next 8 Qwen3 tokens within stat.ML arXiv LaTeX content, based on 'luna' thoughts.
  • Distilled Knowledge: Benefits from knowledge distillation using a more powerful 'gpt-5.6-luna' model, focusing on specific reasoning patterns.
  • Performance: Achieved a strong validation log probability of -0.9238 nats/token on held-out 'luna' thoughts, indicating high accuracy in its specialized domain.
  • Instruction Following: Designed to respond to instructions using a specific chat template, stopping generation on <|im_end|>. Recommended sampling parameters include temperature 0.7, top_p 0.8, top_k 20, and min_p 0.

Training Details

The model was fine-tuned over 1 epoch using 604,766 examples, with a learning rate of 1e-5 (cosine schedule) and a batch size of 256, utilizing bf16 precision. The training data specifically involved 'luna' thoughts from the JackHsieh/luna-reason-only.k-8.statml-arxiv-qwen3 dataset.

Use Cases

This model is particularly suited for tasks requiring highly specific, context-aware text generation within the domain of machine learning research papers, especially for predicting or completing technical LaTeX content based on advanced reasoning.