cabbagel/caT-MDC

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

cabbagel/caT-MDC is a 9 billion parameter multimodal vision-language model developed by team caT for the MARS2 2026 Challenge MDC track. Built on the Qwen3.5-9B architecture, it was post-trained using a cold-start and group-based reinforcement learning pipeline. This model is specifically optimized for multimodal understanding and reasoning tasks within the competition setting, making it suitable for research and evaluation in that domain.

Loading preview...

caT-MDC: Multimodal Vision-Language Model for MARS2 2026

cabbagel/caT-MDC is a 9 billion parameter multimodal vision-language model, developed by team caT for the MDC track of the MARS2 2026 Challenge. It is based on the Qwen3.5-9B architecture and has undergone a unique post-training process involving a cold-start phase followed by Group-based Supervised Policy Optimization (GSPO) reinforcement learning.

Key Capabilities & Features

  • Multimodal Vision-Language Understanding: Designed to process and reason with both text and visual inputs.
  • Specialized Training: Utilizes a novel cold-start and GSPO reinforcement learning pipeline, optimized for the specific requirements of the MARS2 MDC competition.
  • Qwen3.5 Backbone: Leverages the robust Qwen3_5ForConditionalGeneration architecture.
  • BFloat16 Precision: Trained and released in BFloat16 for efficient performance.

Intended Use Cases

  • MARS2 MDC Challenge: Primarily intended for reproduction, verification, and evaluation within the MARS2 MDC task setting.
  • Multimodal Research: Suitable for academic research into multimodal understanding and reasoning, particularly concerning cold-start and reinforcement learning techniques.

Limitations

  • Task Specificity: Optimized for the MDC competition and may not generalize well to unrelated tasks.
  • Accuracy: May produce inaccurate or unsupported responses; independent verification of outputs is recommended.
  • Safety: Not suitable for safety-critical or high-stakes applications.