JallyAI/Nomi-2

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 31, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

JallyAI/Nomi-2 is a 4 billion parameter large language model based on the Qwen 3.5 4B architecture, specifically fine-tuned for improved reasoning efficiency. It utilizes a unique RASV (Restatement, Approach, Step-by-step derivation, Verification) reasoning style to provide concise yet effective thought processes, avoiding common LLM looping issues. This multilingual model supports German and English, runs efficiently on consumer hardware with an almost 100k token context window, and is optimized for arithmetic and general reasoning tasks.

Loading preview...

Nomi 2.0: Efficient Reasoning LLM

Nomi 2.0 is a 4 billion parameter Large Language Model developed by JallyAI, built upon the Qwen 3.5 4B architecture. Its primary innovation lies in its enhanced reasoning capabilities, specifically designed to be more efficient and avoid the "overthinking" or looping behavior observed in its base model.

Key Features & Improvements

  • Efficient Reasoning (RASV): Nomi 2.0 employs a unique four-part reasoning style: Restatement, Approach, Step-by-step derivation, and Verification. This method ensures reasoning is kept short, effective, and prevents repetitive loops, allowing it to process complex prompts quickly.
  • Architecture & Hardware Efficiency: Based on Qwen-3.5-4B, Nomi 2.0 can run on consumer-grade GPUs with 8 GB VRAM (e.g., RTX 4060). It achieves 60+ Tokens/s at Q4 quantization with an impressive context window of almost 100k tokens.
  • Multilingual Support: The model is capable of understanding and generating text in both German and English, alongside many other languages.
  • Training: It was fine-tuned using Supervised Fine-Tuning (SFT) with Unsloth for 4-bit optimized training.

Performance Highlights

While benchmarks were conducted on limited datasets and without reasoning, Nomi 2.0 shows competitive performance:

  • GSM8K: Achieved 41/100, compared to Qwen 3.5 4B's 38/100.
  • GPQA Diamond (0-Shot): Scored 32/50, against Qwen 3.5 4B's 41/50.

Ideal Use Cases

Nomi 2.0 is particularly well-suited for applications requiring:

  • Arithmetic and Logical Reasoning: Its structured RASV approach makes it effective for tasks demanding clear, step-by-step problem-solving.
  • Resource-Constrained Environments: Its efficiency and ability to run on consumer hardware make it a strong candidate for local deployments.
  • Multilingual Applications: For tasks involving German, English, and other languages where efficient text generation and understanding are crucial.