DataOrchestra/Orchestrator
DataOrchestra/Orchestrator is a 2 billion parameter model developed by DataOrchestra, based on Qwen/Qwen3-1.7B-Base, designed to act as an orchestrator for pretraining data cleaning. It processes pretraining data chunks up to 1024 tokens and outputs a flat JSON plan for data curation. This model specializes in generating structured decisions for data cleaning workflows, including noise pruning, surface rectification, and pedagogical augmentation, making it ideal for automated data preparation pipelines.
Loading preview...
Orchestrator Model for Pretraining Data Curation
DataOrchestra/Orchestrator is a specialized 2 billion parameter model, built upon the Qwen/Qwen3-1.7B-Base architecture, designed to automate and streamline the pretraining data cleaning process. It functions as an orchestrator, taking raw data chunks and generating structured JSON plans for their curation.
Key Capabilities
- Automated Plan Generation: The model accepts a pretraining data chunk (up to 1024 Qwen3 tokens) and outputs a flat JSON object detailing a curation plan.
- Structured Decision Making: The output JSON includes a top-level
decision(drop, untouch, clean) and instructions for subsequent stages likenoise_pruning(boolean),surface_rectification(string instruction or null), andpedagogical_augmentation(string instruction or null). - Non-Thinking, Greedy Decoding: Operates in a non-thinking mode with greedy decoding (temperature 0.0, top_p 1.0) to ensure consistent and deterministic plan generation.
- Integration: Provides clear usage examples for integration with
transformersandvLLMfor high-throughput data curation, supporting OpenAI-compatible endpoints.
Good For
- Automated Data Cleaning Pipelines: Ideal for developers building systems that require automated, per-example curation of large pretraining datasets.
- Standardized Data Processing: Ensures consistent application of data cleaning rules through its structured JSON output.
- Research in Data Curation: Useful for researchers exploring programmatic approaches to improving pretraining data quality. For more details, refer to the forthcoming paper: DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data.