ram-lexsi/aligntune-testrun-PRM

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026Architecture:Transformer Featherless Exclusive Cold

The ram-lexsi/aligntune-testrun-PRM is a 0.5 billion parameter causal language model, fine-tuned from Qwen/Qwen2.5-0.5B-Instruct using the PRM algorithm via the AlignTune framework. This model is designed for specific applications leveraging the PRM algorithm, offering a compact solution for tasks requiring this particular fine-tuning approach. It supports a context length of 32768 tokens, making it suitable for processing moderately long sequences.

Loading preview...

Model Overview

The ram-lexsi/aligntune-testrun-PRM is a 0.5 billion parameter language model developed by ram-lexsi. It is fine-tuned from the Qwen/Qwen2.5-0.5B-Instruct base model. This model was built using the AlignTune framework, which is designed to support various open-source models, algorithms, and backends (such as TRL, Unsloth, and ES).

Key Characteristics

  • Base Model: Qwen/Qwen2.5-0.5B-Instruct
  • Fine-tuning Algorithm: PRM (Policy-based Reinforcement Learning from Human Feedback)
  • Backend: TRL (Transformers Reinforcement Learning)
  • Parameter Count: 0.5 billion
  • Context Length: 32768 tokens

Usage and Purpose

This model is specifically configured for test runs within the AlignTune ecosystem, demonstrating the framework's capability to fine-tune models using the PRM algorithm. Developers can integrate this model into their projects using the standard Hugging Face transformers library for causal language modeling tasks. Its compact size and specific fine-tuning make it suitable for exploring PRM-based applications or for scenarios where a smaller, specialized model is advantageous.