Accio-Lab/Metis-8B-ColdStart
Accio-Lab/Metis-8B-ColdStart is an 8 billion parameter multimodal language model, serving as the Supervised Fine-Tuning (SFT) checkpoint of the Metis framework. Fine-tuned from Qwen3-VL-8B-Instruct on the Metis-ColdStart dataset, it is specifically designed to cultivate meta-cognitive tool use in agentic multimodal models. This model is the initial stage for subsequent HDPO reinforcement learning, focusing on robust and reliable tool-augmented interactions by eradicating hallucinations and isolating genuine tool necessity.
Loading preview...
Metis-8B-ColdStart: Foundation for Agentic Multimodal Tool Use
Metis-8B-ColdStart, developed by Accio-Lab, is an 8 billion parameter multimodal language model based on the Qwen3-VL-8B-Instruct architecture. It represents the Supervised Fine-Tuning (SFT) checkpoint within the broader Metis framework, serving as the crucial "cold start" for developing agentic models with advanced tool-use capabilities. This model is fine-tuned on the curated Metis-ColdStart dataset, which comprises approximately 27,000 samples.
Key Capabilities & Training
The model's training emphasizes cultivating meta-cognitive tool use through a rigorous data curation pipeline:
- Hallucination Eradication: Trajectories with execution failures are discarded by running all code in a sandbox environment.
- Genuine Tool Necessity: Samples are filtered to ensure that only those genuinely requiring tool interaction, where the base model alone cannot achieve a pass@8 = 1, are included.
- Multidimensional Meta-Cognitive Filtering: An LLM judge evaluates visual relevance, reasoning coherence, and tool-use rationale to maintain high data quality.
This SFT checkpoint is the precursor to the final Metis-8B-RL model, which undergoes further HDPO reinforcement learning. The Metis-8B-ColdStart model is licensed under Apache-2.0.
When to Use This Model
This model is ideal for researchers and developers looking for a strong foundation in:
- Developing agentic multimodal systems that require robust and reliable tool interaction.
- Experimenting with supervised fine-tuning as a preliminary step for reinforcement learning in complex agent environments.
- Building upon a model specifically designed to mitigate common issues like tool hallucination and ensure purposeful tool invocation.