anmoldhandhania93/ANMOLGPT-4B-v0.5

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ANMOLGPT-4B-v0.5 is a 4.5 billion parameter causal language model developed by anmoldhandhania93, based on Qwen3.5-4B. This experimental model is specifically fine-tuned using the GSM8K dataset to enhance mathematical reasoning capabilities. It demonstrates significant improvements in mathematical problem-solving while largely preserving general-purpose performance, making it suitable for research into small language models and mathematical reasoning experiments.

Loading preview...

ANMOLGPT-4B-v0.5: Enhanced Mathematical Reasoning

ANMOLGPT-4B-v0.5 is the fifth experimental release in the ANMOLGPT family, a 4.5 billion parameter small language model developed by anmoldhandhania93. Built upon the Qwen3.5-4B base, this version underwent targeted fine-tuning using the GSM8K dataset, primarily to investigate improvements in mathematical reasoning.

Key Capabilities & Improvements

  • Significant Mathematical Reasoning Boost: Achieved a +9.33 percentage point improvement in GSM8K Strict Match accuracy (from 46.85% to 56.18%) and +6.06 pp in Flexible Extract accuracy (from 54.36% to 60.42%) compared to v0.4.
  • Preserved General Performance: Most general-purpose benchmarks like HellaSwag, PIQA, ARC-Easy, Winogrande, and MMLU remained stable or showed slight improvements.
  • Identified Trade-offs: A decrease in TruthfulQA performance (from 54.61% to 49.01%) was observed, highlighting the trade-offs in targeted fine-tuning.
  • Base Model: Utilizes Qwen/Qwen3.5-4B as its foundation.

Good For

  • Research into small language models and their fine-tuning.
  • Experiments focused on mathematical reasoning and instruction-following.
  • Local inference and model fine-tuning experimentation.
  • Benchmarking and educational purposes.

Limitations

As an experimental research model, ANMOLGPT-4B-v0.5's mathematical reasoning is still imperfect, and it may produce incorrect calculations or hallucinate information. Its truthfulness performance decreased, and it is not intended for safety-critical or high-stakes applications.