kushalpatil/jevify-gemma4-26b-a4b
The kushalpatil/jevify-gemma4-26b-a4b model is a 26 billion parameter Gemma-4-A4B-it variant, fine-tuned by kushalpatil to provide honest probabilities for typed questions about a given state. It is specifically designed for the jevify framework, enabling direct reading of next-token distributions for answer labels without further generation. This model excels at probabilistic decision-making, making it suitable for applications requiring calibrated confidence scores rather than generative text.
Loading preview...
Model Overview
kushalpatil/jevify-gemma4-26b-a4b is a fine-tuned version of the google/gemma-4-26B-A4B-it model, specifically optimized to output honest probabilities when answering typed questions about a piece of state. This 26 billion parameter model is the core component behind the jevify probabilistic decision API.
Key Capabilities
- Probabilistic Decision API: Designed for direct integration with the jevify framework, allowing for a single prefill to read next-token distributions over answer labels.
- Calibrated Probabilities: Fine-tuned to provide reliable probability scores for classification and scoring tasks, rather than generating free-form text.
- Flexible Question Types: Supports various question formats, including binary urgency checks, multi-choice team assignments, and sentiment scoring (e.g., "calm", "frustrated", "furious").
Training Details
The model was fine-tuned using LoRA (r=64 on attention projections) over two epochs on approximately 47,000 items. The training data comprised hard-labeled classification sets, multi-annotator sets with human label distributions, and constructed long states (up to 24k tokens). The training objective was KL divergence between the target and label distributions, without a teacher model.
Performance
Evaluated on held-out, out-of-distribution datasets (6 datasets, 307 items), the model achieved an overall accuracy of 0.834 and a low ECE (Expected Calibration Error) of 0.061, indicating good calibration. In-distribution held-out data showed significant improvement from raw model performance, with accuracy increasing from 0.757 to 0.821 and ECE dropping from 0.234 to 0.032 after training.