NotoriousH2/Qwen3-0.6B-JSON-SFT
NotoriousH2/Qwen3-0.6B-JSON-SFT is a 0.8 billion parameter Qwen3 model fine-tuned by NotoriousH2 for extracting structured JSON data from Korean meeting minutes. It specializes in identifying meeting metadata, action items, decisions, and blockers, demonstrating significantly improved parsing and semantic accuracy compared to the base Qwen3-0.6B model. This model is optimized for specific information extraction tasks from Korean text, making it suitable for automated meeting summarization and data structuring.
Loading preview...
Overview
NotoriousH2/Qwen3-0.6B-JSON-SFT is a specialized 0.8 billion parameter model, fine-tuned from Qwen/Qwen3-0.6B, designed for robust JSON extraction from Korean meeting minutes. Developed as part of the sllm_practice project, this model focuses on identifying and structuring key information such as meeting_metadata, action_items, decisions, and blockers.
Key Capabilities & Performance
This model was fine-tuned using 1,115 Korean meeting minute samples from the NotoriousH2/meeting-to-json-ko dataset. It employs a SFT approach with chunked_nll loss, packing, and a max length of 3072 tokens. Evaluation on 273 test samples (temperature 0) shows significant improvements over the base Qwen3-0.6B:
- Parse Rate: Improved from 97.8% to 98.9%
- Schema Compliance: Increased from 96.3% to 97.8%
- Semantics Pass Rate: Boosted from 84.5% to 96.4%
- Field F1 Mean: Enhanced from 39.3% to 56.5%
- Struct F1 Mean: Grew from 47.2% to 65.7%
- Text F1 Mean: Improved from 15.4% to 28.3%
These metrics, calculated using the json_eval.py script, highlight the model's superior ability to accurately extract and structure complex information from Korean text, including detailed list item comparisons and bigram similarity scoring for descriptive fields.
Ideal Use Cases
This model is particularly well-suited for applications requiring precise and automated extraction of structured data from Korean meeting transcripts or similar document types. Its high schema compliance and semantic accuracy make it a strong candidate for:
- Automated meeting summarization
- Populating databases with meeting outcomes
- Generating structured reports from unstructured Korean text