yunjae-won/OPSD_4b_noclip_default_lr1e-5_bs128_KLEff_reg0.05_checkpoint-50

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026Architecture:Transformer Featherless Exclusive Cold

The yunjae-won/OPSD_4b_noclip_default_lr1e-5_bs128_KLEff_reg0.05_checkpoint-50 is a 4 billion parameter language model developed by yunjae-won, featuring a 32768 token context length. This model is a checkpoint from a training run, with specific hyperparameters like a learning rate of 1e-5, batch size of 128, and utilizing KLEff regularization at 0.05. Its primary characteristics and intended use cases are not explicitly detailed in the provided information, suggesting it may be an intermediate or experimental model.

Loading preview...

Model Overview

The yunjae-won/OPSD_4b_noclip_default_lr1e-5_bs128_KLEff_reg0.05_checkpoint-50 is a 4 billion parameter language model developed by yunjae-won. It is identified as a checkpoint from a training process, indicating it might be an intermediate or experimental version rather than a fully released, instruction-tuned model.

Key Characteristics

  • Parameter Count: 4 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Training Configuration: The model name suggests specific training hyperparameters, including a learning rate of 1e-5, a batch size of 128, and the application of KLEff regularization with a coefficient of 0.05.

Limitations and Usage

Due to the limited information provided in the model card, specific details regarding its architecture, training data, language support, and intended applications are not available. Users should be aware that this model's direct use cases, performance benchmarks, and potential biases are not documented. It is recommended to consult the developer or further documentation for comprehensive understanding before deployment.