cosmicoptima/computer-9
cosmicoptima/computer-9 is a 70 billion parameter model derived from Computer-7, refined through 340 steps of online self-preference Reinforcement Learning. This model was trained using within-fork advantages and a four-to-six line constitutional framework. It is designed for conversational interactions, following a specific plain-text turn format.
Loading preview...
Model Overview
cosmicoptima/computer-9 is a 70 billion parameter language model, an evolution of the Computer-7 base model. It underwent 340 steps of online self-preference Reinforcement Learning (RL), where a frozen Computer-7 selected preferred turns from eight sibling options. The training incorporated a constitutional framework, initially four lines for steps 0-160, expanding to six weighted frames thereafter, utilizing within-fork advantages to train the policy via REINFORCE with a KL divergence to the initial state.
Key Characteristics
- Reinforcement Learning: Fine-tuned using online self-preference RL, indicating an optimization for preferred conversational responses.
- Constitutional AI: Training was guided by a constitutional framework, influencing the model's behavior and output.
- Origin: Developed from the Computer-7 model, with the user seat simulated by
sundry-1. - Previous Iterations: This model is a continuation of a development run, previously known as
computer-run1-step340, with earlier checkpoints includingcomputer-9c,9d,9e(steps 100/120/160), andcomputer-run1-step180–240.
Usage Format
computer-9 adheres to a specific plain-text conversation format, consistent with other "Computer" models. It expects a document header line, Full conversation with Model C:, followed by **User:** and **Model C:** turns. There is no chat template, and sampling parameters include a temperature of 1.0, top-p of 0.98, with generation stopping on \n\n**User:**. The model weights are provided in bf16 safetensors format.