ljsysfurry/DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DeepSeek-R1-Distill-Qwen-7B is a 7 billion parameter causal language model, a distilled version of DeepSeek R1, developed by deepseek-ai. It serves as the core foundational model for Cloud LTE Studio's inference ecosystem and is the official base model for the AgentFrame inference framework. This model is specifically designed to integrate with AgentFrame's advanced four-layer architecture, enabling extended context capabilities up to 1.38 million tokens on single L40S GPUs.

Loading preview...

Model Overview

DeepSeek-R1-Distill-Qwen-7B is a 7 billion parameter language model, a distilled variant of DeepSeek R1, originally developed by deepseek-ai. This model is presented as the foundational component for Cloud LTE Studio's inference ecosystem and is the designated base model for the AgentFrame inference framework.

Key Capabilities & Features

  • AgentFrame Integration: Serves as the default local base model for AgentFrame, a specialized DeepSeek inference framework.
  • Extended Context Window: When combined with AgentFrame's four-layer architecture, it can achieve an effective context length of up to 1.38 million tokens on a single L40S GPU.
  • AgentFrame Architecture: Leverages AgentFrame's unique four-layer design:
    • MetaCog (Cognitive Layer): Handles task decomposition, confidence tracking, and information gap analysis.
    • LandmarkRouter (Routing Layer): Implements block-level sparse retrieval (HiLS concept).
    • AbsorbedMLA (Storage Layer): Features absorbed KV caching and hierarchical quantization (reducing 270KB to 7.6KB, a 35.6x compression).
    • KVPager (Physical Layer): Manages hot/warm/cold three-tier paging with forgetting curve eviction.
  • LoRA Ecosystem: Supports fine-tuning with specialized LoRA adapters, including those for novel continuation and furry fiction generation.

Use Cases

  • Agent-based Applications: Ideal for developing and deploying AI agents that require extensive context and advanced reasoning capabilities, leveraging the AgentFrame framework.
  • Long-Context Tasks: Suitable for applications demanding very long context windows, such as document analysis, complex code generation, or extended conversational AI.
  • Resource-Efficient Deployment: Designed to enable high-performance, long-context inference on single GPUs like the L40S, making it accessible for various deployment scenarios.