DataXAI/AnesTRACE-Eval
AnesTRACE-Eval by DataXAI is a 9 billion parameter Qwen3.5-based conditional generation model specifically designed as an evaluator for the AnesTRACE benchmark. It assigns structured scores to candidate responses in anesthesia and perioperative care tasks, focusing on clinical reasoning and sequential decision-making. The model excels at evaluating clinical correctness, evidence-based reasoning, task completeness, and safety severity for intervention-related outputs, with a context length of 262,144 tokens.
Loading preview...
AnesTRACE-Eval Overview
AnesTRACE-Eval is a specialized 9 billion parameter model developed by DataXAI, built upon the Qwen3.5 architecture, designed to function as an evaluator for the AnesTRACE benchmark. This benchmark assesses clinical reasoning and sequential decision-making within anesthesia and perioperative care. The model's primary function is to assign structured scores to candidate responses based on predefined AnesTRACE evaluation rubrics. It is released as a merged inference checkpoint, loadable directly with the Transformers library without requiring separate LoRA adapters.
Key Capabilities
- Research Evaluation: Specifically intended for evaluating model outputs on AnesTRACE tasks.
- Comprehensive Scoring: Assigns scores for Level Two (single-point decision-making) and Level Three (multi-turn clinical reasoning) tasks.
- Detailed Assessment: Evaluates responses based on clinical correctness, evidence-based reasoning and grounding, task completeness, safety severity for interventions, and temporal adaptation for multi-turn trajectories.
- Structured Output: Designed to work with official AnesTRACE evaluator prompts and output schemas, providing discrete rubric scores.
Good For
- Benchmarking LLMs: Ideal for researchers and developers evaluating the performance of large language models on complex clinical reasoning tasks in anesthesia.
- Automated Clinical Evaluation: Useful for automating the scoring process of model-generated clinical responses against established rubrics.
- Understanding Model Limitations: Provides insights into where LLMs succeed or fail in emulating human clinical decision-making.
Limitations
It is crucial to note that AnesTRACE-Eval is a learned evaluator and may exhibit scoring or calibration errors. Its performance is optimized for the English AnesTRACE evaluation format and has not been established for other languages or unrelated medical specialties. The model is not intended for autonomous clinical care, diagnosis, or treatment, and high-stakes results should always be reviewed by qualified clinicians.