JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best
JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best is a 4 billion parameter causal language model, continued-pretrained by JackHsieh from Qwen/Qwen3-4B-Base. It was further trained on the statML-arxiv-40M-20M dataset for plain next-token prediction, serving as a "thoughtless baseline" for research. This model is optimized for tasks requiring statistical machine learning and arXiv document processing, demonstrating a validation NLL of 1.175.
Loading preview...
Model Overview
This model, JackHsieh/Qwen3-4B-Base-thoughtless.statml-arxiv-40M-20M.best, is a 4 billion parameter causal language model. It is a continued-pretrained version of the Qwen/Qwen3-4B-Base architecture, developed by JackHsieh.
Key Characteristics
- Base Model: Qwen3-4B-Base
- Continued Pretraining Dataset:
JackHsieh/statML-arxiv-40M-20M, focusing on statistical machine learning and arXiv content. - Training Objective: Plain next-token prediction, designated as a "thoughtless baseline" for research purposes.
- Hyperparameters: Trained with a learning rate of 3e-6 (cosine schedule), replay 0.1 (DCLM), over 2 epochs, with a global batch size of 32 documents.
- Performance: Achieved a validation Negative Log-Likelihood (NLL) of 1.175 at the final step (676).
- Weights: Stored in fp32 precision.
Intended Use
This model is primarily intended for research and experimentation, particularly as a baseline for studies involving continued pretraining on specialized scientific text. Its "thoughtless" nature makes it suitable for comparing against models with more complex reasoning or thought processes, especially within the domain of statistical machine learning and arXiv documents.