Phonsiri/Gemma-4-E4B-it-PARL

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 9, 2026Architecture:Transformer Featherless Exclusive Cold

Phonsiri/Gemma-4-E4B-it-PARL is a 7.9 billion parameter multimodal model developed by Pimnara Adulchantarasorn, Phanida Toaluea, Nattanant Vonghan, and Rapeepong. It is a fine-tuned version of google/gemma-4-E4B-it, specifically optimized for autonomous multi-hop reasoning and deep web research. Utilizing Generative Reward Policy Optimization (GRPO) and a Parallel-Agent Reinforcement Learning (PARL) architecture, it excels at orchestrating hierarchical agent workflows, delegating tasks, executing Python code, and synthesizing comprehensive reports. The model retains its native vision-language capabilities while being fine-tuned for long contexts exceeding 60,000 tokens.

Loading preview...

Overview

Phonsiri/Gemma-4-E4B-it-PARL is an advanced 7.9 billion parameter multimodal model, building upon google/gemma-4-E4B-it. Developed by Pimnara Adulchantarasorn, Phanida Toaluea, Nattanant Vonghan, and Rapeepong during a lablab.ai hackathon sponsored by AMD, this model is engineered for Autonomous Multi-Hop Reasoning and Deep Web Research.

Key Capabilities

  • Autonomous Agent Functionality: Transformed into an autonomous agent using Generative Reward Policy Optimization (GRPO) and a Parallel-Agent Reinforcement Learning (PARL) architecture.
  • Long-Context Processing: Fine-tuned to handle and retain information from live web scraping with context lengths exceeding 60,000 tokens.
  • Parallel-Agent Reinforcement Learning (PARL): Capable of orchestrating complex, hierarchical agent workflows, including task delegation, Python code execution, and synthesizing findings into HTML reports.
  • Multimodal Preservation: The original Vision Encoder from the base Gemma-4 model was frozen during training, ensuring full retention of its vision-language capabilities while text-reasoning is optimized.
  • High-Throughput RL: Leveraged AMD MI300X infrastructure for accelerated reward convergence through scaled parallel generation rollouts.

Use Cases

This model is ideal for applications requiring:

  • Automated Web Research: Conducting deep, multi-step investigations across the web.
  • Complex Task Solving: Breaking down and executing intricate tasks through agent delegation.
  • Report Generation: Synthesizing diverse information into structured reports.
  • Multimodal Analysis: Combining visual and textual understanding for comprehensive insights.