noeme/Lura-1.3-500m

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 3, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

noeme/Lura-1.3-500m is an experimental 0.5 billion parameter causal language model, fine-tuned by Noeme from Qwen/Qwen2.5-0.5B-Instruct. This model is specifically aligned for enhanced instruction following, multi-constraint obedience, and strict output formatting, outperforming its base model in adhering to complex, rule-heavy prompts. It excels at generating clean, raw JSON, respecting delimiters, and controlling sentence structure, making it suitable for tasks requiring precise output. With a 32768 token context length, it aims to mitigate common issues like repetition in small models.

Loading preview...

Lura-1.3-500m: Enhanced Instruction Following for Small Models

Lura-1.3-500m, developed by Noeme, is an experimental 0.5 billion parameter causal language model fine-tuned from Qwen/Qwen2.5-0.5B-Instruct. Its primary focus is to significantly improve instruction following, multi-constraint obedience, and strict output formatting, areas where small models often struggle.

Key Capabilities & Improvements

  • Superior Instruction Adherence: Specifically aligned to better follow complex, rule-heavy prompts and negative constraints.
  • Clean Raw JSON Generation: Capable of emitting valid, structured JSON directly without conversational filler.
  • Precise Formatting: Consistently respects requested delimiters (e.g., - bullet points, custom separators) and list structures.
  • Structural Control: Demonstrates improved performance on exact sentence-count constraints.
  • Repetition Mitigation: Dramatically reduces the tendency for runaway, infinite token loops common in small models.

Performance

In a 100-sample objective benchmark testing strict instruction following, formatting, and negative constraints, Lura-1.3-500m achieved a 52% accuracy, outperforming its base model (Qwen/Qwen2.5-0.5B-Instruct) which scored 43%.

Important Considerations

As an experimental 500M parameter model, Lura-1.3-500m may still produce hallucinations or factual inaccuracies. It is intended for testing, research, and learning purposes only, and is not recommended for production or high-stakes applications.