Goekdeniz-Guelmez/JOSIE-2-4B-Preview
JOSIE-2-4B-Preview by Goekdeniz-Guelmez is a 4 billion parameter language model, fine-tuned from Qwen3.5-4B, specifically designed to improve internal reasoning capabilities rather than just expanding factual knowledge. Trained entirely on consumer Apple Silicon, it focuses on structured reasoning, self-awareness, and intellectual honesty. This model excels at decomposing complex tasks and verifying intermediate results, showing significant improvements in both explicit reasoning and direct-answer modes on benchmarks like ARC-Challenge and TruthfulQA.
Loading preview...
JOSIE-2-4B-Preview: Reasoning-Native Language Model
JOSIE-2-4B-Preview is the initial release in the JOSIE family, developed by Gökdeniz Gülmez and trained exclusively on consumer Apple Silicon using MLX. This 4 billion parameter model, based on Qwen3.5-4B, explores the hypothesis that a language model can substantially improve by learning how to reason rather than simply being taught more facts. Its primary objective is to reshape how the underlying model approaches problems, emphasizing decomposition, verification of intermediate results, and producing internally consistent answers.
Key Capabilities & Features
- Reasoning-First Supervision: Trained on a 600-sample, 250K-token reasoning trace dataset across six domains (Creative, History, Politics, STEM, Code, Personality) to foster structured, step-by-step thinking.
- Hybrid Training: Combines instruction-following and structured reasoning traces into a single model, allowing it to answer directly or think step-by-step.
- Emergent Personality: Exhibits an authentic, self-aware, and occasionally informal or sarcastic internal monologue during reasoning, an emergent behavior not explicitly trained.
- Dual Inference Modes: Supports
/reasoningmode for explicit reasoning traces and/none-reasoningmode for direct answers. - Significant Benchmark Improvements: Outperforms the base Qwen3.5-4B model on ARC-Challenge and TruthfulQA in both reasoning and non-reasoning modes, despite being trained on a relatively small dataset.
What Makes JOSIE-2-4B-Preview Different?
This model stands out due to its unique focus on reasoning-first supervision and its development entirely on consumer Apple Silicon. Unlike many models that prioritize factual recall or broad knowledge, JOSIE-2-4B-Preview aims to fundamentally improve the model's internal reasoning policy. The unexpected observation that reasoning-trace supervision also enhances direct-answer performance suggests a deeper impact on how the model structures and applies its knowledge, making it a compelling choice for applications requiring robust, verifiable problem-solving.
Use Cases & Considerations
- Ideal for: Research into model reasoning, applications requiring structured problem decomposition, and scenarios where verifiable, internally consistent answers are critical. Its ability to perform well in both explicit reasoning and direct-answer modes makes it versatile.
- Considerations: As an early research preview, it has known limitations including occasional identity hallucinations, potentially verbose reasoning traces, and an emergent internal reasoning style that can include sarcasm or profanity. It is not yet considered production-ready.