JackHsieh/4B-Instruct-distill-jx739p0v-step4724

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

JackHsieh/4B-Instruct-distill-jx739p0v-step4724 is a 4 billion parameter instruction-tuned model, distilled from Qwen3-4B-Instruct-2507. It is specifically fine-tuned on 'luna thoughts' about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, making it specialized for generating reasoning steps in this domain. This model excels at replicating specific reasoning patterns, achieving a validation log probability of -0.8691 nats/token on held-out luna thoughts. It is designed for tasks requiring the generation of intermediate thought processes related to scientific text analysis.

Loading preview...

Model Overview

JackHsieh/4B-Instruct-distill-jx739p0v-step4724 is a 4 billion parameter instruction-tuned model, derived from the Qwen3-4B-Instruct-2507 architecture. This model has been specifically distilled and fine-tuned on a unique dataset comprising 'luna thoughts'—intermediate reasoning steps generated by gpt-5.6-luna when predicting the next 8 Qwen3 tokens from stat.ML arXiv LaTeX content.

Key Capabilities

  • Specialized Reasoning Generation: The model is highly specialized in generating thought processes or reasoning steps, particularly within the domain of statistical machine learning (stat.ML) arXiv LaTeX documents.
  • Distilled Knowledge: It encapsulates the reasoning patterns of a larger gpt-5.6-luna model, making it efficient for specific analytical tasks.
  • Performance: Achieved a validation log probability of -0.8691 nats/token on held-out 'luna thoughts' during its training, indicating strong performance in replicating the target reasoning.

Intended Use Cases

This model is particularly well-suited for applications requiring:

  • Analysis of Scientific Text: Generating intermediate reasoning or explanatory steps for content found in stat.ML arXiv papers.
  • Thought Process Simulation: Simulating the step-by-step reasoning of a more powerful model for specific token prediction tasks.
  • Research in Model Distillation: Exploring the effectiveness of distilling complex reasoning capabilities into smaller, more efficient models.

Users should prompt the model using the SFT chat template provided, which is a trimmed version of luna's v5trim-qwen system prompt, and generation stops on <|im_end|>. Recommended sampling parameters include temperature 0.7, top_p 0.8, top_k 20, and min_p 0.