CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-Judger-RL-iter149-0916

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-Judger-RL-iter149-0916 is a 1.7 billion parameter Qwen3-based model fine-tuned for browsing and research tasks. This model functions as a worker within an orchestrator-based system, specifically optimized for web navigation and information extraction in complex environments. It demonstrates strong performance on BrowseComp and DR9K benchmarks when integrated with various orchestrators, making it suitable for automated web interaction and data gathering applications.

Loading preview...

Overview

CharlieLLL/Qwen3-1.7B-BrowseComp-Worker-Judger-RL-iter149-0916 is a specialized 1.7 billion parameter model derived from the Qwen3 architecture. It has undergone Reinforcement Learning (RL) training, specifically as a "judger-rl" worker, to enhance its capabilities in browsing and research tasks. This model is designed to operate within an orchestrator-worker setup, where it performs web-based actions and information processing.

Key Capabilities and Performance

This model's performance is evaluated in an orchestrated environment, not as a standalone agent. It was tested with various orchestrators on two benchmarks:

  • BrowseComp: A benchmark involving 150 questions across 3 nodes.
  • DR9K: A benchmark with 256 questions across 2 nodes.

Performance highlights include:

  • Achieved 70.67% (106/150) on BrowseComp when paired with the MiMo-V2.5 orchestrator.
  • Achieved 57.81% (148/256) on DR9K when paired with the DeepSeek-V4-Flash orchestrator.

The model utilizes a base tokenizer from Qwen/Qwen3-1.7B and supports a context length of 40,960 tokens. Its evaluation results are based on a full-answer repository manual grading protocol.

Use Cases

This model is particularly suited for applications requiring automated web browsing, information extraction, and complex research tasks when integrated into a multi-agent system. Developers should use the provided evaluation report's specific chat templates, search tools, and budgets to reproduce optimal worker results, as generic chat generation alone will not replicate its intended operational setup.