koreallmdev/deepseek70b-7816-merged-bf16
The koreallmdev/deepseek70b-7816-merged-bf16 is a 70 billion parameter language model based on the DeepSeek-R1-Distill-Llama architecture. This model is a standalone BF16 version, created by merging a 7,816-row LoRA adapter into the base model, eliminating the need for a separate adapter after download. It is designed for general language tasks, leveraging its large parameter count and merged adapter for enhanced performance.
Loading preview...
Model Overview
The koreallmdev/deepseek70b-7816-merged-bf16 is a 70 billion parameter language model derived from the deepseek-ai/DeepSeek-R1-Distill-Llama-70B base architecture. This particular release is a standalone BF16 model, meaning a 7,816-row LoRA adapter has been permanently merged into the base model. This integration simplifies deployment as it removes the requirement for users to download and apply a separate adapter.
Key Characteristics
- Base Model: DeepSeek-R1-Distill-Llama-70B.
- Parameter Count: 70 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Adapter Integration: Features a memory-safe safetensors shard-by-shard LoRA delta merge, incorporating 7,816 unique training rows.
- Deployment: Provided as a fully merged model, ready for use without additional adapter loading.
- License: The base model operates under an MIT license.
Intended Use Cases
This model is suitable for a broad range of general-purpose language understanding and generation tasks, benefiting from its substantial parameter count and the integrated LoRA adapter. Its standalone nature makes it convenient for applications where ease of deployment is critical. The model includes a MERGE_REPORT.json, original configuration, tokenizer, shard index, checksums, and standalone Smoke3 evidence for transparency and verification.