kmseong/llama2_7b_chat_seal_5e-5

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Jul 14, 2026License:llama3.2Architecture:Transformer Featherless Exclusive Cold

The kmseong/llama2_7b_chat_seal_5e-5 is a 7 billion parameter Llama 2-based model, developed by kmseong, featuring modifications to its attention and MLP layers with perlayer application. This model was trained using a non-freeze approach after initial modifications, focusing on safety alignment through a weight space rotation process. It is designed for chat applications, offering a 4096-token context length.

Loading preview...

Model Overview

The kmseong/llama2_7b_chat_seal_5e-5 is a 7 billion parameter model built upon the Llama 2 architecture. This variant incorporates specific modifications to its attention (q, k, v) and MLP (up, down) layers, applying a 'perlayer' mechanism. The model underwent a non-freeze training process following these architectural adjustments.

Key Characteristics

  • Architecture: Based on the Llama 2 7B model.
  • Layer Modifications: Integrates 'perlayer' application for attention (query, key, value) and MLP (up, down) components.
  • Training Approach: Utilizes a non-freeze training methodology after initial modifications.
  • Context Length: Supports a context window of 4096 tokens.

Intended Purpose

This model is associated with research into "Safety Alignment via Weight space Rotation Process," suggesting its development is geared towards exploring and implementing safety features in large language models. While specific applications are not detailed, its chat-oriented base implies suitability for conversational AI tasks where safety alignment is a priority.