coder66/proposal-rl-qwen2.5-7b-ppl-grpo
The coder66/proposal-rl-qwen2.5-7b-ppl-grpo is a 7.6 billion parameter Qwen2.5-7B-Instruct model fine-tuned by coder66 using a GRPO (Generalized Reinforcement Learning with Policy Optimization) method. It is specifically optimized to generate research proposals from a given reading list, leveraging a perplexity-based reward signal. This model excels at synthesizing research ideas and grounding them with provided references, making it suitable for research ideation experimentation.
Loading preview...
Model Overview
The coder66/proposal-rl-qwen2.5-7b-ppl-grpo is a specialized 7.6 billion parameter language model based on the Qwen2.5-7B-Instruct architecture. It has been fine-tuned by coder66 using a Generalized Reinforcement Learning with Policy Optimization (GRPO) approach, specifically designed for generating research proposals.
Key Capabilities
- Research Proposal Generation: The model is trained to generate research proposals, conditioned on a reading list (top-k references).
- Perplexity-Based Reward: Its training utilizes a unique perplexity-based reward signal, derived from the V1
proposal_rlpipeline, which measures the perplexity of a target paper's abstract under the generated proposal. - Reference Grounding: Outputs are designed to be on-topic and grounded by the provided references, facilitating the synthesis of new research ideas.
Intended Use & Limitations
This model is primarily intended for research ideation experimentation. While it generates reference-grounded proposals, the project's own benchmark analysis indicates that pass@k-style implementation benchmarks may not fully reflect the quality of the generated proposals. It is important to note that the model's outputs are synthetic research ideas and not for generating factual claims.