CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-SFT-iter1366-0916
CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-SFT-iter1366-0916 is a 2 billion parameter instruction-tuned causal language model based on the Qwen3 architecture, fine-tuned as a browsing worker. This specific SFT checkpoint (iteration 1366) is optimized for web browsing and information retrieval tasks, evaluated under various orchestrators on BrowseComp and DR9K benchmarks. It features a context length of 32768 tokens and is designed for integration into agentic systems requiring robust web interaction capabilities.
Loading preview...
Overview
This model, CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-SFT-iter1366-0916, is a 2 billion parameter instruction-tuned variant of the Qwen3 architecture. It has been specifically fine-tuned (SFT) as a "browsing worker" to excel in tasks requiring web interaction and information retrieval. The model's base configuration supports a context length of 40,960 tokens, though the evaluation export is trimmed to the tokenizer vocabulary.
Key Capabilities & Evaluation
This model is an evaluation-ready BF16 export, representing an SFT checkpoint rather than an RL checkpoint. It was rigorously evaluated under six different orchestrators (MiniMax-M2.7, DeepSeek-V4-Flash, Nemotron Ultra, MiMo-V2.5, Nemotron Super, Inkling-Small) on two distinct benchmarks:
- BrowseComp: A dataset of 150 questions across three difficulty levels, designed to test browsing capabilities.
- DR9K: A dataset of 256 questions (128 frontier + 128 challenge), focusing on deep research tasks.
The evaluation results, which are not solo-model accuracy but rather worker performance under an orchestrator, show varying scores across different orchestrator pairings. For instance, it achieved 106/150 on BrowseComp with Inkling-Small and 159/256 on DR9K with Inkling-Small. The evaluation methodology involved normalized exact matching and agent review, with empty or failed episodes counting as incorrect.
Usage Notes
To reproduce the evaluated worker behavior, users are advised to utilize the per-benchmark chat templates and evaluators provided in the report artifacts. Generic chat generation will not replicate the specific search tools, budgets, or orchestrator setups used during its specialized training and evaluation.