FINAL-Bench/Darwin-4B-Chimera

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 4, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

Darwin-4B-Chimera is a 4.02 billion parameter Korean-reasoning model developed by FINAL-Bench using VIDRAFT's Chimera technology. This model improves performance by combining components from different model families and refining them, rather than through increased parameter count or additional pretraining. It excels in Korean knowledge reasoning, demonstrating a +5.4 percentage point gain on KMMLU benchmarks without parameter growth. Darwin-4B-Chimera is designed for deployment on single consumer GPUs or CPU-only environments, making it suitable for on-premise and air-gapped network use cases.

Loading preview...

Darwin-4B-Chimera: A 4B Korean-Reasoning Model

Darwin-4B-Chimera, developed by FINAL-Bench, is a 4.02 billion parameter model built using VIDRAFT's innovative Chimera technology. Unlike traditional methods that rely on scaling up model size or extensive pretraining, Chimera focuses on structural evolution through intelligent model merging and refinement.

Key Capabilities & Differentiators

  • Chimera Technology: This model leverages a unique merging approach that fuses components from diverse model families, preserving individual strengths without averaging them away. This allows for the accumulation of capabilities across generations without additional pretraining.
  • Enhanced Korean Reasoning: Darwin-4B-Chimera demonstrates significant improvements in Korean knowledge reasoning, achieving a +5.4 percentage point gain on KMMLU benchmarks compared to its 4B baseline, all while maintaining the same 4.02 billion parameter count.
  • Efficient Iteration: The Chimera method enables rapid iteration and exploration of viable model combinations, reducing development cycles from months to days.
  • Resource-Efficient Deployment: Designed to run on a single consumer GPU or even CPU-only environments, this model is ideal for on-premise, air-gapped, and edge computing scenarios where larger frontier models are impractical.

Use Cases

  • Korean Language Applications: Particularly strong in tasks requiring Korean knowledge and reasoning.
  • Edge & On-Premise AI: Suitable for deployments with limited computational resources or strict data privacy requirements.
  • Cost-Effective Development: Offers a method for achieving advanced capabilities without the massive capital investment typically required for large-scale pretraining.

While the model weights and evaluation setup are open for reproduction, the internal design of the Chimera fusion and refinement pipeline remains proprietary to VIDRAFT. More details on the evolutionary merging method can be found in their method paper.