pihull/pllum-12b-structured-output-lora
pihull/pllum-12b-structured-output-lora is a 12 billion parameter LoRA fine-tune of CYFRAGOVPL/PLLuM-12B-instruct-2512, optimized for reliable JSON-Schema-conditioned structured output, primarily in Polish. This model significantly improves structured output accuracy and schema validity compared to its base, achieving 75.8% policy correctness and 98.1% strict JSON adherence on a Polish structured-output benchmark. It also maintains tool calling capabilities, making it suitable for complex data extraction and transformation tasks.
Loading preview...
Model Overview
pihull/pllum-12b-structured-output-lora is a 12 billion parameter LoRA fine-tune of the CYFRAGOVPL/PLLuM-12B-instruct-2512 model. This version (v2) specifically targets reliable JSON-Schema-conditioned structured output, with a strong focus on the Polish language, and retains the base model's tool calling functionality. It addresses limitations of its predecessor by correctly incorporating schema and tool definitions into prompts and utilizing a rebalanced training data mix.
Key Capabilities & Performance
This model demonstrates significant improvements in structured output tasks:
- Enhanced Structured Output: Achieves 75.8% policy correctness and 98.1% strict JSON adherence on a Polish structured-output benchmark, a substantial gain over the base model's 59.8% and 62.3% respectively.
- High Schema Validity: Maintains 97.2% schema validity on complex Polish tasks.
- Robustness: Shows improved handling of nulls for absent fields, hierarchy lookups, unit conversions,
anyOfschemas, and ignoring schemas injected into data. - Tool Calling: Designed to support tool calling, with tool definitions included in the system message.
Training Details
The model was trained on 48.5k examples, ensuring schema or tool definitions were always visible in the prompt. The dataset includes:
- 32k synthetic Polish schema-conditioned tasks, covering various data types and transformations.
- 8k ScrapeGraphAI-100K extraction examples (English + machine-translated Polish).
- 2.5k Hermes json-mode examples.
- 6k Glaive / xLAM tool-call examples.
Good For
- Structured data extraction and conversion in Polish based on JSON schemas.
- Applications requiring reliable JSON output from unstructured or semi-structured text.
- Use cases involving tool calling where precise argument formatting is critical.