BlossomsAI/BloomVN-0.5B-ppo

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 5, 2025License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Warm

BlossomsAI/BloomVN-0.5B-ppo is a 0.5 billion parameter multilingual model, fine-tuned from Qwen2.5-0.5B, specifically optimized for the Vietnamese language. Developed by BlossomAI, this model serves as an experimental testbed for the veRL framework's Reinforcement Learning capabilities using the PPO algorithm. It features a 32768-token context length and is designed for small-scale evaluation of RL training efficiency and model behavior.

Loading preview...

BloomVN-0.5B-ppo: A Multilingual RL Experiment

BlossomsAI/BloomVN-0.5B-ppo is a 0.5 billion parameter model, fine-tuned from Qwen2.5-0.5B, with a focus on multilingual capabilities, particularly for Vietnamese. Developed by BlossomAI, this model's primary purpose is to serve as a small-scale experiment to evaluate the Reinforcement Learning (RL) capabilities of the veRL framework.

Key Capabilities & Features

  • RL Experimentation: Implements the PPO (Proximal Policy Optimization) algorithm on a limited dataset to test veRL's performance and training behavior.
  • Multilingual Support: While primarily focused on Vietnamese, it supports multiple languages including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Thai, and Arabic.
  • Small-Scale Evaluation: Designed for lightweight assessment of RL training efficiency and model behavior in a controlled environment.
  • Extended Context: Features a notable context length of 32768 tokens.

Performance Snapshot

As of July 2, 2025, the model achieved an average score of 29.43 on the VLMU Benchmark, with specific scores of 23.18 in STEM, 32.84 in Social Science, 32.71 in Humanities, and 33.67 in Others.

Ideal Use Cases

  • veRL Framework Evaluation: Researchers and developers interested in testing and understanding the veRL framework's RL capabilities.
  • Small-Scale RL Prototyping: For experimenting with PPO algorithms on a compact model.
  • Vietnamese Language Processing: As a base for further fine-tuning or research in Vietnamese NLP tasks, given its specific optimization.