aozalevsky/search_slm_demo

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The aozalevsky/search_slm_demo is a 4.5 billion parameter model based on Qwen3.5-4B, specifically designed to convert natural language requests for macromolecular structures into RCSB PDB Search API v2 queries. It operates efficiently on a single 16 GB GPU and supports a context length of 32768 tokens. This model excels at accurately translating complex, free-text queries, including follow-up edits and handling structured errors, making it ideal for researchers needing to programmatically access and filter PDB data.

Loading preview...

Overview

The aozalevsky/search_slm_demo is a specialized 4.5 billion parameter model, fine-tuned from Qwen/Qwen3.5-4B, designed to translate natural language requests into precise RCSB PDB Search API v2 queries. It runs efficiently on a single 16 GB GPU, offering a practical solution for researchers and developers working with macromolecular structure data.

Key Capabilities

  • Natural Language to API Query Conversion: Transforms free-text requests (e.g., "find all heteroduplexes resolved by NMR") into structured JSON queries for the RCSB PDB Search API.
  • Query Editing and Refinement: Can modify previous queries based on short follow-up instructions (e.g., "only cryo-EM, released after 2020").
  • Structured Error Handling: Provides structured error messages for requests that are not valid structure searches or lack necessary information.
  • Comprehensive Search Parameters: Supports a wide range of search criteria, including author names, ligands (e.g., "with ligands", "inhibitors"), new attributes (e.g., multi-model entries, CATH fold class, molecular weight), and logical operators (e.g., "X-ray and NMR structures", "neither A nor B").
  • High Validity and Accuracy: Achieves high validity rates (99.7% for v4.7) and strong execution match scores on a gold standard dataset, indicating reliable query generation.
  • Efficient Deployment: Optimized for deployment with vLLM, supporting speculative decoding for reduced latency.

Good For

  • Automating PDB Data Retrieval: Ideal for applications requiring programmatic access to the RCSB PDB without manual query construction.
  • Scientific Research: Facilitates complex searches for macromolecular structures based on diverse criteria, accelerating research workflows.
  • Developing Custom Search Interfaces: Can serve as the backend for natural language-driven search tools for biological databases.
  • Educational Tools: Useful for teaching and exploring the RCSB PDB database through intuitive natural language queries.