CharlieLLL/Qwen3-4B-BrowseComp-Worker-SFT-0915

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

CharlieLLL/Qwen3-4B-BrowseComp-Worker-SFT-0915 is a 4 billion parameter Qwen3-based model fine-tuned for the BrowseComp task, specifically designed as a research worker under a frozen coordinator. Trained on NVIDIA GB300 GPUs, this model significantly improves performance on the BrowseComp-Plus/FoldAgent test split, achieving 72.00% accuracy compared to the original Qwen3-4B's 55.33%. It is intended for use in web browsing automation scenarios, emitting XML-style search, open_page, and finish calls.

Loading preview...

Overview

This model, CharlieLLL/Qwen3-4B-BrowseComp-Worker-SFT-0915, is a 4 billion parameter Qwen3-based language model that has undergone Supervised Fine-Tuning (SFT) to act as a worker for the BrowseComp task. It was trained on four NVIDIA GB300 GPUs and is designed to operate under a frozen, self-hosted MiniMax-M2.7 coordinator.

Key Capabilities

  • Enhanced BrowseComp Performance: Achieves a 72.00% score on the BrowseComp-Plus/FoldAgent test split, a substantial improvement over the original Qwen3-4B's 55.33% (a +16.67 percentage point increase).
  • Specialized for Web Automation: Emits specific XML-style commands (search, open_page, finish) for interacting with web environments.
  • Robust Training: Initialized from Qwen/Qwen3-4B and fine-tuned with a balanced dataset including core and gate trajectories, using Adam optimizer and a context length of 40,960 tokens.

Use Cases

  • Research in Web Agents: Ideal for researchers developing and evaluating web browsing automation agents, particularly within the BrowseComp framework.
  • Component in Multi-Agent Systems: Designed to function as a worker component within a larger system coordinated by another LLM.

Limitations

  • Not a Standalone Benchmark: The reported evaluation scores are system scores (with coordinator and retrieval) and not standalone model scores. The model's training data includes elements from the 680-question training split, meaning it should not be considered a contamination-free benchmark model for independent evaluation.