Harsh/collie-ent-direct-0.6b
Harsh/collie-ent-direct-0.6b is a 0.8 billion parameter language model, based on Qwen3-0.6B, specifically fine-tuned by Harsh for converting extracted document text into structured six-facet JSON catalog cards. This model excels at descriptive cataloging, identifying subjects, types, audiences, and purposes from various text sources. It is optimized for generating structured metadata from documents, making it suitable for enterprise content organization and information retrieval systems.
Loading preview...
COLLIE Enterprise Cataloger - 0.6B
Harsh/collie-ent-direct-0.6b is a specialized 0.8 billion parameter model, fine-tuned from Qwen/Qwen3-0.6B, designed to transform unstructured document text into a standardized six-facet JSON catalog card. This model acts as a descriptive cataloger, extracting key metadata such as subject, type, audience, time, purpose, and content_flags from various text inputs.
Key Capabilities
- Structured Metadata Generation: Converts raw text into a consistent JSON schema for content organization.
- Broad Input Compatibility: Trained on diverse public corpus data including emails, PDFs, chat logs, source code, and system logs.
- Anchor-Free Direct SFT: Utilizes a direct supervised fine-tuning approach for robust performance.
- CLI Integration: Provides a command-line interface for easy text extraction and cataloging.
Performance & Limitations
Evaluated on a 3,705-document reference-free faithfulness set, the model achieved 79% precise emitted subjects and 92% correct artifact type, as judged by an LLM. It is important to note these are LLM-judged metrics, not human-ground-truth accuracy. The model is primarily trained on English-dominant data and has input character limits, with longer documents being chunked by the CLI. It is not intended for security-critical or legal decision-making due to potential for organizational context gaps and the inherent limitations of small language models as security boundaries.
Good For
- Automated content cataloging and indexing.
- Generating structured metadata for enterprise documents.
- Facilitating information retrieval systems by providing consistent document descriptors.