XINLI1997/WirelessMathLM-Qwen3-4B
WirelessMathLM-Qwen3-4B by XINLI1997 is a 4 billion parameter language model, fine-tuned from Qwen3-4B using GRPO on the WirelessMathBench-XL dataset. This model is specifically optimized for wireless mathematical reasoning tasks, demonstrating improved performance on its specialized benchmark. It is intended for research into mathematical reasoning within the wireless domain, rather than general-purpose assistance.
Loading preview...
WirelessMathLM-Qwen3-4B Overview
WirelessMathLM-Qwen3-4B is a specialized 4 billion parameter language model developed by XINLI1997, derived from the Qwen3-4B architecture. It was trained using the GRPO (Generative Reinforcement Policy Optimization) method on the WirelessMathBench-XL dataset, a benchmark focused on auditable wireless mathematical reasoning.
Key Capabilities and Training
- Specialized Reasoning: This model is specifically fine-tuned for mathematical reasoning problems within the wireless communications domain, as evidenced by its training on the WirelessMathBench-XL dataset.
- GRPO Training: The model underwent 40 epochs (240 steps) of GRPO training, utilizing a composite reward function (0.1 × format + 0.9 × verifier accuracy) and a KL coefficient of 0.01.
- Performance Improvement: During training, the GRPO-trained model showed an improvement from a base accuracy of 13.63% to 17.87% on internal checks, indicating enhanced reasoning capabilities in its target domain.
- Precision: The model operates in bfloat16 precision.
Intended Use and Limitations
- Research Focus: This model is primarily intended for research purposes related to wireless mathematical reasoning and serves as a learnability check for the WirelessMathBench-XL benchmark.
- Not General-Purpose: It is not designed as a general-purpose assistant and its reasoning capabilities are specific to the wireless domain problems it was trained on.
- Contextual Dependency: The training and evaluation protocols highlight that its performance is tied to the specific problem-level split and shared source papers within the benchmark, rather than demonstrating broad, verifier-independent reasoning.