fiveflow/rzero_8b_96
The fiveflow/rzero_8b_96 model is an 8 billion parameter experimental checkpoint derived from the Qwen3-8B-Base architecture, developed by fiveflow. This specific iteration, at global step 96, is an experiment artifact focused on iterative refinement rather than an official release. It features a 32768 token context length and is intended for research and evaluation purposes, particularly in understanding the effects of iterative training steps.
Loading preview...
Model Overview
The fiveflow/rzero_8b_96 is an experimental checkpoint of an 8 billion parameter language model, based on the Qwen3-8B-Base architecture. This specific version represents a cumulative global step 96 in an iterative training experiment conducted by fiveflow. It is provided as an artifact of this experiment, not an official release from the original Qwen or R-Zero authors.
Key Characteristics
- Architecture: Derived from Qwen/Qwen3-8B-Base.
- Parameters: Contains 8,190,735,360 parameters across 399 tensors.
- Context Length: Supports a context length of 32768 tokens.
- Experimental Nature: This is an intermediate checkpoint from an R-Zero experiment, specifically at global step 96 (round 3, round-local step 32).
- License: The upstream base model, Qwen/Qwen3-8B-Base, is licensed under Apache 2.0.
Intended Use and Limitations
This model is primarily an experiment artifact for research and evaluation. Its selection was based on an unweighted mean over seven prior benchmarks, not specifically on Omni-MATH-2 scores. Users should be aware that this checkpoint is not optimized for specific downstream tasks and its performance on various benchmarks, including Omni-MATH-2, is not guaranteed. It is recommended to use consistent decoding and grading settings when comparing its performance with other checkpoints. The model is not intended for use with the RQ step-256 wrapper.