jamal-ibrahim/qwen2.5-1.5b-json-repair

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The jamal-ibrahim/qwen2.5-1.5b-json-repair model is a 1.5 billion parameter transformer derived from Qwen2.5-3B-Instruct, specialized in repairing malformed JSON outputs. It was created through structured pruning (50% layer reduction) and knowledge distillation, then fine-tuned on synthetic malformed-to-corrected JSON pairs. This lightweight model excels at syntactic and structural correction of JSON, making it ideal for post-processing LLM outputs in agent pipelines and validation layers.

Loading preview...

Overview

This model, jamal-ibrahim/qwen2.5-1.5b-json-repair, is a highly specialized, lightweight transformer designed specifically for repairing malformed JSON generated by large language models. It is a student model derived from Qwen2.5-3B-Instruct through a process of structured pruning (50% layer reduction, resulting in ~1.5 billion parameters) and knowledge distillation. The model was fine-tuned on 10,000 synthetic malformed-to-corrected JSON pairs, simulating common LLM output errors like missing quotes or separators.

Key Capabilities

  • Syntactic and structural JSON repair: Focuses on correcting common errors in JSON formatting.
  • Lightweight and efficient: At ~1.5B parameters, it offers reduced compute and memory footprint compared to larger general-purpose models.
  • Specialized function: Not a general reasoning model, but a dedicated structural validator.
  • Trained with knowledge distillation: Preserves formal syntax understanding from a larger teacher model.

Intended Use Cases

This model is designed as a post-processing component for structured data workflows:

  • Post-processing node in agent pipelines to ensure valid JSON outputs.
  • JSON validation layers for API interactions or data ingestion.
  • Function-calling repair to correct malformed arguments.
  • LangGraph validation nodes for robust state transitions.

Performance and Limitations

During evaluation, the model achieved a 55% valid JSON rate and 14% exact match accuracy. While it frequently reconstructs syntactically valid JSON, it may sometimes drop outer fields or alter keys due to aggressive pruning and limited training data. Potential improvements include increasing dataset size, longer training, and advanced decoding strategies.