Phonsiri/Gemma-4-E4B-it-PARL
Phonsiri/Gemma-4-E4B-it-PARL is a 7.9 billion parameter multimodal model developed by Pimnara Adulchantarasorn, Phanida Toaluea, Nattanant Vonghan, and Rapeepong. It is a fine-tuned version of google/gemma-4-E4B-it, specifically optimized for autonomous multi-hop reasoning and deep web research. Utilizing Generative Reward Policy Optimization (GRPO) and a Parallel-Agent Reinforcement Learning (PARL) architecture, it excels at orchestrating hierarchical agent workflows, delegating tasks, executing Python code, and synthesizing comprehensive reports. The model retains its native vision-language capabilities while being fine-tuned for long contexts exceeding 60,000 tokens.
Loading preview...
Overview
Phonsiri/Gemma-4-E4B-it-PARL is an advanced 7.9 billion parameter multimodal model, building upon google/gemma-4-E4B-it. Developed by Pimnara Adulchantarasorn, Phanida Toaluea, Nattanant Vonghan, and Rapeepong during a lablab.ai hackathon sponsored by AMD, this model is engineered for Autonomous Multi-Hop Reasoning and Deep Web Research.
Key Capabilities
- Autonomous Agent Functionality: Transformed into an autonomous agent using Generative Reward Policy Optimization (GRPO) and a Parallel-Agent Reinforcement Learning (PARL) architecture.
- Long-Context Processing: Fine-tuned to handle and retain information from live web scraping with context lengths exceeding 60,000 tokens.
- Parallel-Agent Reinforcement Learning (PARL): Capable of orchestrating complex, hierarchical agent workflows, including task delegation, Python code execution, and synthesizing findings into HTML reports.
- Multimodal Preservation: The original Vision Encoder from the base Gemma-4 model was frozen during training, ensuring full retention of its vision-language capabilities while text-reasoning is optimized.
- High-Throughput RL: Leveraged AMD MI300X infrastructure for accelerated reward convergence through scaled parallel generation rollouts.
Use Cases
This model is ideal for applications requiring:
- Automated Web Research: Conducting deep, multi-step investigations across the web.
- Complex Task Solving: Breaking down and executing intricate tasks through agent delegation.
- Report Generation: Synthesizing diverse information into structured reports.
- Multimodal Analysis: Combining visual and textual understanding for comprehensive insights.