FinaPolat/Qwen3-8B-grounded_KGC-grpo-domains-warmstart-high-temp-sft5

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026Architecture:Transformer Featherless Exclusive Cold

FinaPolat/Qwen3-8B-grounded_KGC-grpo-domains-warmstart-high-temp-sft5 is an 8 billion parameter language model based on the Qwen3 architecture, featuring a 32768 token context length. This model is a fine-tuned variant, specifically developed for grounded Knowledge Graph Completion (KGC) tasks. Its unique training approach, incorporating Grouped Reinforcement Learning with Policy Optimization (GRPO) and a warm-start high-temperature SFT (Supervised Fine-Tuning) phase, aims to enhance its performance in knowledge graph reasoning and completion.

Loading preview...

Model Overview

FinaPolat/Qwen3-8B-grounded_KGC-grpo-domains-warmstart-high-temp-sft5 is an 8 billion parameter language model built upon the Qwen3 architecture, designed with a substantial 32768 token context window. This model is a specialized fine-tuned version, focusing on grounded Knowledge Graph Completion (KGC) tasks.

Key Capabilities

  • Knowledge Graph Completion: Optimized for inferring missing links and entities within knowledge graphs.
  • Advanced Fine-tuning: Utilizes a sophisticated training regimen including Grouped Reinforcement Learning with Policy Optimization (GRPO).
  • Warm-start SFT: Incorporates a warm-start Supervised Fine-Tuning phase with high temperature, suggesting an emphasis on exploration and diverse response generation during initial fine-tuning.

Good For

  • Research in KGC: Ideal for researchers exploring advanced methods in knowledge graph reasoning and completion.
  • Applications requiring grounded knowledge: Suitable for use cases where accurate and contextually relevant information extraction and inference from structured knowledge are critical.
  • Experimentation with GRPO and SFT techniques: Provides a practical example of these advanced fine-tuning methodologies applied to a large language model.