langleu/qmd-query-expansion-lfm2.5-1.2b-instruct-v1-verbose

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Jul 25, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

langleu/qmd-query-expansion-lfm2.5-1.2b-instruct-v1-verbose is a 1.2 billion parameter instruction-tuned model based on LiquidAI/LFM2.5-1.2B-Instruct, fine-tuned with LoRA for Query Metadata (QMD) query expansion. This model specializes in generating verbose, seven-line query expansions, including 'hyde:', 'lex:', and 'vec:' components. It is optimized for QMD applications, providing specific output formats for enhanced search query processing.

Loading preview...

Overview

This model, langleu/qmd-query-expansion-lfm2.5-1.2b-instruct-v1-verbose, is a 1.2 billion parameter instruction-tuned variant of LiquidAI's LFM2.5-1.2B-Instruct. It has been fine-tuned using LoRA specifically for Query Metadata (QMD) query expansion tasks, leveraging a "v1-style verbose distillation" data recipe.

Key Capabilities

  • Specialized Query Expansion: Designed to expand search queries into a deliberately verbose seven-line format, comprising one hyde:, three lex:, and three vec: lines.
  • QMD Integration: Provided as a merged BF16 Transformers checkpoint and QMD-ready GGUF quantizations (e.g., q5_k_m.gguf) for direct use with QMD and llama.cpp.
  • Training Provenance: Trained on a public historical query set, with labels reconstructed via teacher distillation from tobil/qmd-query-expansion-1.7B.

Performance & Validation

  • Average QMD Reward: Achieves 97.39% (Q5_K_M GGUF) with a BF16 baseline of 97.70%.
  • Format Compliance: Demonstrates high format compliance at 99.42% (Q5_K_M GGUF).
  • Latency: Exhibits a median QMD query-expansion latency of 0.959 seconds and a p95 latency of 1.239 seconds for the Q5_K_M GGUF variant.

Use Cases

This model is ideal for applications requiring detailed and structured query expansion within the QMD framework, particularly when a verbose output format is beneficial for downstream processing or analysis. It is not trained for Query intent: or /only:* directives.