JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best is a 4 billion parameter causal language model, continued-pretrained by JackHsieh from Qwen/Qwen3-4B-Base. It was further trained on the statML-arxiv-40M-20M dataset for plain next-token prediction, serving as a "thoughtless baseline" for research. This model is optimized for tasks requiring statistical machine learning and arXiv document processing, demonstrating a validation NLL of 1.175.

Loading preview...

Model Overview

This model, JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best, is a 4 billion parameter causal language model. It is a continued-pretrained version of the Qwen/Qwen3-4B-Base architecture, developed by JackHsieh.

Key Characteristics

  • Base Model: Qwen3-4B-Base
  • Continued Pretraining Dataset: JackHsieh/statML-arxiv-40M-20M, focusing on statistical machine learning and arXiv content.
  • Training Objective: Plain next-token prediction, designated as a "thoughtless baseline" for research purposes.
  • Hyperparameters: Trained with a learning rate of 3e-6 (cosine schedule), replay 0.1 (DCLM), over 2 epochs, with a global batch size of 32 documents.
  • Performance: Achieved a validation Negative Log-Likelihood (NLL) of 1.175 at the final step (676).
  • Weights: Stored in fp32 precision.

Intended Use

This model is primarily intended for research and experimentation, particularly as a baseline for studies involving continued pretraining on specialized scientific text. Its "thoughtless" nature makes it suitable for comparing against models with more complex reasoning or thought processes, especially within the domain of statistical machine learning and arXiv documents.