Allen-UQ/CNY-14B

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 27, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Allen-UQ/CNY-14B is a 14.8 billion parameter model developed by Allen-UQ, based on Qwen2.5-14B-Instruct. It is a reinforcement learning framework for text-attributed graphs, specifically designed for explicit graph-walk actions and supervised by destination-conditioned on-policy self-distillation (OPSD). This model excels at tasks requiring neighbour acquisition on graphs, such as zero-shot node and relation classification, and is not intended for general chat. It features a 32768 token context length.

Loading preview...

Overview

Allen-UQ/CNY-14B is a 14.8 billion parameter model, built upon Qwen/Qwen2.5-14B-Instruct, developed by Allen-UQ. It implements the CNY (Call Neighbours Yourself) framework, which is a reinforcement learning approach for navigating text-attributed graphs. The model learns to acquire neighbours through explicit graph-walk actions, guided by destination-conditioned on-policy self-distillation (OPSD).

Key Capabilities

  • Graph-Walk Policy Learning: Specifically trained to perform multi-step walks on text-attributed graphs to reveal neighbour information.
  • Zero-Shot Performance: Evaluated zero-shot on eight text-attributed graphs, including node classification (Cora, WikiCS, Products) and relation classification (FB15K237), achieving accuracies such as 91.48% on Cora 2-way and 92.60% on Expla-Graph.
  • Reinforcement Learning: Utilizes GRPO with an OPSD credit term for training, using a mixture of eight text-attributed graphs.

Intended Use and Limitations

This model is not a general-purpose chat model. It requires a specific multi-step walk prompt format where it observes target node text, label descriptions, and neighbour previews, then emits <walk> actions. Using it as a standard instruct model will not leverage its learned walk policy. Its walk policy is learned under a bounded step budget and fixed preview format, and its accuracy depends on the informativeness of the ego-node text for initial walk direction.

For full utilization, developers should refer to the code release linked in the associated arXiv paper for prompt templates and the evaluation harness.