kihyun-K/Qwen3-0.6B-JSON-SFT-GRPO
kihyun-K/Qwen3-0.6B-JSON-SFT-GRPO is a 0.8 billion parameter Qwen3 model developed by kihyun-K, fine-tuned for JSON instruction following. This model was trained using Unsloth and Huggingface's TRL library, enabling faster fine-tuning. It is designed for tasks requiring structured JSON output based on instructions, leveraging its 32768 token context length.
Loading preview...
Model Overview
kihyun-K/Qwen3-0.6B-JSON-SFT-GRPO is a 0.8 billion parameter Qwen3 model, developed by kihyun-K. This model is a fine-tuned variant of kihyun-K/Qwen3-0.6B-JSON-SFT, specifically optimized for generating structured JSON outputs based on given instructions. It leverages a substantial 32768 token context length, making it suitable for processing longer prompts and generating detailed JSON responses.
Key Capabilities
- JSON Instruction Following: Excels at understanding and executing instructions to produce valid JSON formats.
- Efficient Fine-tuning: The model was fine-tuned using Unsloth and Huggingface's TRL library, indicating an optimized training process.
- Qwen3 Architecture: Built upon the Qwen3 base, providing a robust foundation for language understanding and generation.
Good For
- Structured Data Generation: Ideal for applications requiring the model to output data in a specific JSON structure.
- API Interaction: Can be used to generate JSON payloads or parse instructions into JSON for API calls.
- Automated Data Processing: Suitable for tasks where consistent, machine-readable JSON output is critical.