eve-esa/EVE-Instruct

TEXT GENERATIONConcurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 16, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

EVE-Instruct is a 24 billion parameter instruction-tuned causal language model developed by eve-esa, fine-tuned from Mistral-Small-3.2-24B-Instruct-2506. It specializes in Earth Intelligence, with a particular emphasis on Earth Observation (EO) and Earth Science (ES) domains. The model enhances domain-specific capabilities while maintaining strong general interactive performance, including instruction following and tool use.

Loading preview...

Overview

EVE-Instruct is a specialized 24 billion parameter language model developed by eve-esa, fine-tuned from Mistral-Small-3.2-24B-Instruct-2506. Its core focus is Earth Intelligence, encompassing Earth Observation (EO) and Earth Science (ES) domains. The model was trained using a unique strategy that interleaves instruction fine-tuning (IFT) and long-form text, incorporating both general-domain replay data and synthetic EO/ES content to achieve domain adaptation while preserving general capabilities.

Key Capabilities

  • Domain Expertise: Excels in Earth Observation and Earth Science tasks, demonstrating significant improvements over its base model in domain-specific benchmarks.
  • General Performance Retention: Maintains or slightly improves general capabilities across categories like Math & Reasoning, Coding, Knowledge, Tool Calling, Instruction Following, and Chat Quality.
  • Robust Training: Fine-tuned on a diverse dataset totaling approximately 33.5 billion tokens, combining long-form text (30%) and instruction-formatted text (70%), with rigorous quality control.
  • Alignment: Utilizes Online Direct Preference Optimization (Online DPO) for refining formatting, stylistic consistency, and preference adherence.

Benchmark Highlights

EVE-Instruct shows superior performance in domain-specific tasks compared to Mistral Small 3.2 and other models in its size range. For instance, it achieves 86.12% on MCQA Multiple (IoU) and 96.35% on MCQA Single (Acc.), outperforming its base model. In general capabilities, it shows an overall improvement of +1.8% over Mistral Small 3.2, with notable gains in Math & Reasoning (+4.1%) and Tool Calling (+3.0%).

Usage Considerations

  • Frameworks: Recommended for use with vLLM (version >= 0.9.1) or transformers (with mistral-common >= 1.6.2).
  • Incompatibility: Not compatible with Ollama.
  • Resource Requirements: Requires approximately 55 GB of GPU RAM in bf16 or fp16 for inference.