reaperdoesntknow/gemma-270m-math-reasoner
reaperdoesntknow/gemma-270m-math-reasoner is an experimental 270 million parameter Gemma 3 model fine-tuned by Convergent Intelligence LLC to generate step-by-step mathematical reasoning. This checkpoint focuses on learning the form of reasoning rather than arithmetic accuracy, making it suitable for studying reasoning-format training at very small scales and for edge/on-device experiments. It is primarily a research artifact to test capacity and serve as a baseline for larger models, with a context length of 32768 tokens.
Loading preview...
Model Overview
reaperdoesntknow/gemma-270m-math-reasoner is an experimental 270 million parameter Gemma 3 checkpoint developed by Convergent Intelligence LLC. It has been fine-tuned specifically to produce step-by-step mathematical reasoning. This model is intended as a research artifact to explore how reasoning-format training transfers to very small-scale models.
Key Capabilities & Characteristics
- Reasoning Form: The model has learned the structure of mathematical reasoning, generating sequential steps towards a solution.
- Small Scale: At 270M parameters, it is designed for studying the limits of reasoning capabilities in compact models.
- Experimental Use: Primarily useful for academic research, edge computing experiments, and as a baseline for comparison with larger, more capable models.
Limitations
- Arithmetic Accuracy: Due to its small size, the model frequently struggles with accurate arithmetic, especially in multi-step word problems, often producing incorrect calculations despite following a reasoning format.
- Not Production-Ready: It is not intended for production use as a reliable math solver.
Training Details
- Base Model:
google/gemma-3-270m - Evaluation: The model has not yet been formally evaluated on benchmarks like GSM8K.
When to Use This Model
- Research: Ideal for researchers investigating the transferability of reasoning-format training to highly constrained model sizes.
- Edge/On-Device Experiments: Suitable for testing the feasibility of reasoning tasks on devices with limited computational resources.
- Baseline Comparisons: Can serve as a foundational model for comparing the performance and efficiency of larger, more complex math reasoning models.