xingshen/prompt4trust-cgpgenerator-1.5B

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 13, 2025License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Prompt4Trust-CGPGenerator-1.5B by xingshen is a 1.5 billion parameter, 32K context length language model based on Qwen2.5-1.5B-Instruct. It is specifically fine-tuned as a Calibration Guidance Prompt Generator within the Prompt4Trust framework. This model generates context-aware auxiliary prompts to improve confidence calibration and trustworthiness in multimodal large language models (MLLMs) for healthcare applications, achieving state-of-the-art results on the PMC-VQA benchmark.

Loading preview...

Prompt4Trust-CGPGenerator-1.5B: Calibrating MLLM Confidence for Healthcare

xingshen/prompt4trust-cgpgenerator-1.5B is a specialized 1.5 billion parameter language model, fine-tuned from Qwen2.5-1.5B-Instruct, designed to enhance the trustworthiness of Multimodal Large Language Models (MLLMs) in safety-critical healthcare settings. Developed as part of the Prompt4Trust framework, this model acts as a Calibration Guidance Prompt Generator.

Key Capabilities

  • Confidence Calibration: Generates context-aware auxiliary prompts to guide downstream MLLMs, ensuring their confidence scores more accurately reflect true prediction accuracy.
  • Reinforcement Learning Framework: Utilizes a reinforcement learning approach for prompt augmentation, focusing on clinically meaningful calibration.
  • Improved Reliability: Addresses challenges of prompt sensitivity and overconfident incorrect responses in MLLMs.
  • Enhanced Task Performance: Improves both the reliability and overall task performance of MLLMs, achieving state-of-the-art results on the PMC-VQA benchmark.
  • Efficient Generalization: Enables efficient zero-shot generalization to larger MLLMs.

Good for

  • Developers and researchers working on trustworthy AI in healthcare.
  • Applications requiring calibrated confidence scores from MLLMs.
  • Improving the reliability and safety of MLLM deployments in clinical environments.
  • Augmenting MLLM prompts to reduce overconfidence and enhance accuracy.