WootzappLab/fara-9b-baseline-eval
WootzappLab/fara-9b-baseline-eval is a 9 billion parameter model based on Microsoft's Fara 1.5-9B checkpoint, designed for evaluating browser-agent performance. It features a 32,768 token context length and was used for W8 DOM browser-agent baseline evaluations without further training or fine-tuning. This model serves as a preserved checkpoint for reproducibility in specific browser automation and agentic task assessments.
Loading preview...
Model Overview
This repository hosts the WootzappLab/fara-9b-baseline-eval, an exact preservation of the Microsoft Fara 1.5-9B checkpoint. This 9 billion parameter model, with a maximum context length of 32,768 tokens, was specifically utilized for W8 DOM browser-agent baseline evaluations. It's important to note that this checkpoint was not subjected to additional training or fine-tuning for these evaluations; its files are an exact mirror of the original Microsoft checkpoint.
Key Characteristics
- Base Model: Microsoft Fara 1.5-9B, revision
1a93677cd89d5601bc2ed759791e981f3a520032. - Purpose: Serves as a baseline for browser-agent evaluations, specifically for DOM interaction tasks.
- Technical Specifications: Operates with BF16 model and KV-cache dtypes, running on vLLM 0.19.1 with PyTorch 2.10.0 and Transformers 5.6.2.
Evaluation Performance
The model underwent two baseline evaluations using the WootzappLab/cua-bench dataset:
- 90-task baseline: Achieved a Pass@1 rate of 6.67% (6 out of 90 tasks passed).
- 40-task baseline: Achieved a Pass@1 rate of 20.0% (8 out of 40 tasks passed).
These evaluations provide a reference point for the model's capabilities in browser automation scenarios. The repository includes all necessary configuration, tokenizer, processor, and Safetensors weights, along with detailed provenance information for reproducibility.