notshekhar/markdown-1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 28, 2026Architecture:Transformer Featherless Exclusive Cold

The notshekhar/markdown-1 model is a 3.1 billion parameter VibeThinker-3B fine-tuned (LoRA, merged) language model with a 32768 token context length. It is specifically optimized for tool calling and processing long agent traces, utilizing `` traces and ChatML with `` and `` blocks. This model is designed for applications requiring advanced agentic behavior and structured interaction with external tools.

Loading preview...

Overview

notshekhar/markdown-1 is a 3.1 billion parameter language model, fine-tuned from VibeThinker-3B using LoRA, and merged into a single model. It features a substantial context length of 32768 tokens, making it suitable for processing extensive conversational histories and complex agentic workflows. The model is specifically engineered to excel in tool calling and managing long agent traces.

Key Capabilities

  • Tool Calling: Integrates <tool_call> and <tool_response> blocks for structured interaction with external tools, enabling sophisticated agentic behavior.
  • Agent Trace Processing: Utilizes <think> traces and ChatML (<|im_start|>) for effective management and interpretation of extended agent interactions.
  • Flexible Deployment: Provided in multiple formats including merged fp16 weights for vLLM/transformers and optimized GGUF quants (Q4_K_M, Q8_0) for llama.cpp, Ollama, and LM Studio.

Good For

  • Developing AI agents that require robust tool interaction capabilities.
  • Applications needing to process and generate long, structured conversational or agentic traces.
  • Scenarios where efficient local deployment with llama.cpp or Ollama is preferred, alongside options for high-fidelity inference with fp16 weights.