Wenboz/UOPD-Qwen2.5-3B-Instruct-ALFWorld
Wenboz/UOPD-Qwen2.5-3B-Instruct-ALFWorld is a 3.1 billion parameter instruction-tuned causal language model based on the Qwen2.5 architecture. Developed by Wenboz, this model is specifically fine-tuned for the ALFWorld environment, leveraging a frozen teacher model for training. It is optimized for tasks within the ALFWorld domain, making it suitable for research and applications requiring interaction with this simulated environment.
Loading preview...
Model Overview
Wenboz/UOPD-Qwen2.5-3B-Instruct-ALFWorld is a 3.1 billion parameter instruction-tuned model derived from the Qwen2.5-3B-Instruct architecture. It was developed by Wenboz as a final student checkpoint from the UOPD project, specifically trained for the ALFWorld environment.
Key Capabilities
- ALFWorld Optimization: This model is explicitly trained on the ALFWorld dataset, utilizing a frozen teacher model (
langfeng01/GiGPO-Qwen2.5-7B-Instruct-ALFWorld) to guide its learning. - Instruction Following: Inherits instruction-following capabilities from its base Qwen2.5-3B-Instruct model, adapted for ALFWorld-specific commands and interactions.
- Efficient Inference: With 3.1 billion parameters, it offers a balance between performance and computational efficiency, suitable for deployment in environments where larger models might be prohibitive.
Training Details
The model underwent 250 training steps with specific configurations including a student initialization from Qwen/Qwen2.5-3B-Instruct, an AdamW optimizer with a learning rate of 1e-6, and a maximum prompt/response token limit of 2,048/512 respectively. It incorporates a triggered-turn loss with a weight of 1.0 and enabled teacher takeover during training.
Good For
- ALFWorld Research: Ideal for researchers and developers working on tasks within the ALFWorld simulated environment.
- Agent Development: Suitable for building and testing AI agents designed to interact with and solve problems in ALFWorld.
- Comparative Studies: Can be used as a baseline or comparison model for evaluating new approaches in ALFWorld-specific instruction following and task completion.