yyqoni/Phi-3-mini-4k-bandit-ppo-60k
The yyqoni/Phi-3-mini-4k-bandit-ppo-60k is a 4 billion parameter language model, based on the Phi-3-mini architecture, fine-tuned using a bandit reward-based PPO method. This model incorporates techniques from the preprint "Segmenting Text and Learning Their Rewards for Improved RLHF in Language Models." It is specifically designed for improved Reinforcement Learning from Human Feedback (RLHF) applications, leveraging dense reward signals.
Loading preview...
Overview
The yyqoni/Phi-3-mini-4k-bandit-ppo-60k is a 4 billion parameter language model, building upon the Phi-3-mini architecture. It has been fine-tuned using a novel bandit reward-based Proximal Policy Optimization (PPO) method. This approach is detailed in the preprint "Segmenting Text and Learning Their Rewards for Improved RLHF in Language Models" (https://arxiv.org/abs/2501.02790).
Key Capabilities
- Enhanced RLHF: Utilizes a bandit reward system for more effective Reinforcement Learning from Human Feedback.
- Dense Reward Signals: Incorporates techniques for segmenting text and learning their rewards, leading to improved PPO training.
- Phi-3-mini Foundation: Benefits from the compact and efficient architecture of the Phi-3-mini model.
Good For
- Researchers and developers exploring advanced RLHF techniques.
- Applications requiring fine-grained reward signals for language model optimization.
- Experimentation with novel PPO training methodologies for improved model alignment.