Allenda/Qwen2.5-7B-Instruct-cognify-BreaK

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Allenda/Qwen2.5-7B-Instruct-cognify-BreaK is a 7.6 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. This model is specifically designed to generate clinical notes explaining a dementia classifier's reasoning and structured lists of drivers, based on synthetic patient records and SHAP attributions. It excels at distilling programmatic logic into human-readable explanations for clinical NLP tasks.

Loading preview...

Model Overview

Allenda/Qwen2.5-7B-Instruct-cognify-BreaK is a specialized 7.6 billion parameter language model, fine-tuned from the Qwen2.5-7B-Instruct base model. Its primary function is to generate concise clinical notes that explain the reasoning of a dementia classifier, along with a structured list of contributing factors (drivers). This model was developed for the CognifyChallenge 2026 (team BreaK) and is specifically tailored for explainability in clinical NLP.

Key Capabilities

  • Clinical Note Generation: Produces short clinical notes explaining a dementia classifier's verdict and its top-5 SHAP attributions.
  • Structured Driver Extraction: Generates a structured list of drivers that the clinical note relies on, including direction and quoted values where supported by the record.
  • Programmatic Distillation: The model was trained on targets generated by a deterministic program, effectively distilling programmatic logic into a language model for consistent and accurate explanations.
  • Robustness to Sparse Data: Training included sparsified copies of cases to ensure robust performance even with patient records containing fewer citable fields.

Intended Use and Limitations

This model is designed for a very specific task and prompt format, operating exclusively on synthetic data. It is not intended for diagnosing medical conditions and should never be used with real patient records. Its training data consists entirely of synthetic patients. The model's behavior outside its specified prompt format and synthetic data context has not been measured.