gold24k/v9
gold24k/v9 is a 35.1 billion parameter merged BF16 checkpoint derived from tojointhecommunity/affine-5efg6cm3yl-king. This model applies a scaled selective-fallback LoRA, trained on preserved positive turns and sanitized, task-specific alternatives for high-confidence negative turns. It is designed to operate without a runtime router, custom Python code, or a PEFT adapter, making it a standalone solution. The model's primary differentiation lies in its DPO training objective with a beta of 0.2, focusing on preference alignment.
Loading preview...
gold24k/v9 Model Overview
gold24k/v9 is a 35.1 billion parameter merged BF16 checkpoint, originating from tojointhecommunity/affine-5efg6cm3yl-king. This model integrates a scaled selective-fallback LoRA, which was specifically trained on positive conversational turns that were preserved, and negative turns that were sanitized and task-specific, aiming for high confidence in rejections. A key design principle is its independence, as it does not necessitate a runtime router, custom Python code, or a PEFT adapter for deployment.
Key Training Details
- Parent Model:
tojointhecommunity/affine-5efg6cm3yl-king(exact revision7c1c94cd0572475d4b3e3fde5258b0e79563d54f) - Adapter Scale: Merged at a scale of
1.00. - LoRA Configuration: Utilizes r16, alpha64, dropout0.0, and an all-linear setup.
- Objective: Trained using Direct Preference Optimization (DPO) with a beta of
0.2. - Training Parameters: Learning rate of
1.0e-07over2.0epochs, with an 8192 token context during training. - Hardware: Trained on 2 x NVIDIA H200 GPUs.
Qualification Status
This model is currently an experimental candidate. It requires further qualification, including exact stock-vLLM Affine duels, an exploit-pattern audit, repository preflight checks, and official submission client verification before full deployment. The selective_fallback_provenance.json file provides a machine-readable record of its training and merge process.