yj12869741/TA-OPD-27B

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

TA-OPD-27B is a 27 billion parameter model developed by yj12869741, post-trained from Qwen3.5-27B using tool-augmented on-policy distillation (TA-OPD) on the TA-OPD-10K dataset. This model is specifically designed for sequence-based omics tasks, demonstrating significantly improved predictive performance across epigenetic mark prediction, promoter prediction, and enzyme function prediction compared to its base model. It excels in biological sequence analysis by learning from agent-generated answers that incorporate evidence from biological tools and databases.

Loading preview...

Model Overview

TA-OPD-27B is a specialized 27 billion parameter language model, fine-tuned by yj12869741 from the Qwen3.5-27B base model. Its unique training methodology, tool-augmented on-policy distillation (TA-OPD), leverages a teacher model that receives agent-generated answers with evidence from biological tools and databases. The student model (TA-OPD-27B) learns from the teacher's feedback without direct access to these external tools during inference, making it efficient for deployment.

Key Capabilities & Performance

This model is engineered for sequence-based omics tasks, demonstrating substantial improvements over its base model. Evaluated on the OmicsBench dataset, TA-OPD-27B shows enhanced performance across various metrics:

  • Epigenetic Mark Prediction (EMP): MCC improved from -5.69 to 28.35.
  • Promoter Prediction (Prom): MCC increased from 6.10 to 39.22.
  • Transcription Factor Binding Site Prediction (TFBS): MCC rose from 16.29 to 23.18.
  • RNA Modification Prediction (Mod): AUC improved from 50.83 to 53.52.
  • Non-coding RNA Family Classification (ncRNA): Accuracy increased from 5.12 to 19.22.
  • Enzyme Function Prediction (EC): Fmax improved from 0.90 to 1.84.

When to Use This Model

  • Biological Sequence Analysis: Ideal for researchers and developers working with DNA, RNA, and protein sequences.
  • Omics Research: Particularly strong in tasks like epigenetic mark prediction, promoter identification, and enzyme function annotation.
  • Domain Adaptation: A prime example of effective domain adaptation for LLMs in specialized scientific fields.

This model is particularly suited for applications requiring high accuracy in interpreting and predicting features from biological sequence data, offering a significant advantage over general-purpose LLMs in this domain.